AI Blog

Tagged: ecosystem

← Back to AI Blog

10 min read

The 782 is the number about you

GTIG reported on 30 September 2026 that exactly 50% of AI-discovered vulnerabilities yield remote code execution against 26% of everything else — but publishes no sample size, and its attribution method selects for the few vendors currently pointing agents at memory-unsafe systems code. The number worth acting on is four sections down: 782 CVEs in agent frameworks and orchestration in eight months, against 97 for frontier models.

10 min read

A dropped subscription looks exactly like a quiet week

OpenAI shipped plugin automations on all plans on 29 September 2026 against MCP Events — a draft with no SEP number, in a repository whose README calls its contents exploratory. The draft gets the webhook hardening right and makes the two envelopes that report absence optional, so a revoked permission, a lost buffer and a genuinely quiet upstream all reach your agent as the same empty stream.

9 min read

A shared context is a shared credential

At DevDay on 29 September 2026 OpenAI paired always-on Dots agents — each with its own cloud computer, browser and thousands of connectors — with ChatGPT Space, where employees, ChatGPT, Codex and those agents work from one shared context. The permission model people will reason about is per-connector OAuth scope. The boundary that decides what happens is who may write into the context, and nobody is enforcing that one.

12 min read

The tamper-proof half did not ship

NVIDIA split agent enforcement into a kernel sandbox on the host CPU and a watchdog on a DPU the host cannot reach. The sandbox is Apache-2.0 on GitHub today; the watchdog has no ship date. The split is not a release accident — the layer far enough away to be tamper-proof is too far away to understand what the agent was trying to do.

10 min read

Amazon opened the back office and closed the storefront in the same week

On 21 September Amazon cut off Meta's Muse agent; two days later it handed outside AI agents its Seller Central APIs. The variable is not the agent — it is whether a delegation exists that the platform can verify, scope and revoke, which the seller side has had for a decade and the buyer side does not have at all.

12 min read

The ad brought its own agent — OpenAI split the conversation instead of the ranking

Everyone predicted a bought ranking; OpenAI bought the conversation instead, and that is the better design — for exactly as long as the two conversations stay apart. Sponsored Agents put an advertiser-operated agent behind a labelled ad slot and keep it out of the assistant's answer. But the separation is a property of the session, and what actually moves between the two lanes is claims, carried by the person, with no field anywhere saying a paid party said it first.

11 min read

Four reads and one write — Google Home MCP gated the half nobody was worried about

Google blocked the thing everyone asked about: an agent connected through Home MCP cannot unlock your door. But four of the five tools are reads, and list_home_history hands a third-party agent a queryable record of motion, presence and door events over any window — with no equivalent gate, because nobody has written down what a sensitive read is. Actuation is bounded, legible and reversible. The read side is none of those.

9 min read

Meta built Muse assuming the injection lands — and priced the rest at $130,000

The per-user VM is the headline and the least interesting layer. Everything load-bearing in Muse sits downstream of a successful prompt injection — brokered credentials the model never sees, a gatekeeper process the agent cannot argue with, kernel-level taint on anything that read your data — and the bounty schedule says so out loud. The residual risk is not exfiltration; it is the harmful action that travels over an approved channel to an approved destination.

13 min read

CPE is a join key, not a score — NIST is putting an agent inside the NVD

NIST presented its AI agent enrichment workflow for the National Vulnerability Database on 17 September, and the open question is not whether the model is accurate. Enrichment produces three fields that fail in three incompatible ways: a wrong CVSS score gets argued about, a wrong CWE degrades analytics, and a wrong CPE returns no rows at all. One of those failures is silent, and the record format has no field in which a machine can say it was not sure.

8 min read

Target selection just became free — 395 organisations, 48 countries, one operator

GreyNoise published a PaperCut campaign that ran hundreds of AI agents in parallel and reached 440 servers at 395 organisations in 48 countries, 11 of them inside the first 26 seconds. The speed is not the finding. The finding is that choosing who to attack now costs the same as choosing one — which deletes the obscurity discount every mid-size security programme has been quietly spending, and puts the least-resourced sector, education, at the front of the list with 204 victims.

9 min read

WebMCP makes your page an API, and the session is the only auth it has

WebMCP lets a page hand an AI agent a list of callable tools, and Chrome is shipping it behind a flag while the W3C community group draft is still moving. The part worth arguing about is not discovery but authority: a registered tool executes as your page's own JavaScript, inside the session the logged-in user already established, so your server sees a request it cannot distinguish from a click. You are publishing an API whose only credential belongs to someone who is not the caller.

8 min read

The coordinator is the requester now — and nobody scoped the grant

Cursor put Projects into beta on 10 September: a coordinator agent that plans, delegates to thousands of subagents, and — the part worth arguing about — watches a Slack channel, a schedule or all your PRs and acts without waiting for a prompt. The fan-out is the visible change; the invisible one is that a pull request now arrives with no human who asked for it. Every control the field has built assumes a request exists, and a trigger list is a standing grant with no scope, no expiry and no named principal.

10 min read

The Agents API sells you the harness — compaction included

OpenAI opened the Agents API in public beta on 10 September, putting the managed Codex harness — sessions, subagent orchestration, recovery and context compaction — behind one API call, with no fee beyond tokens and containers. The compaction step is the part worth arguing about: it is the transformation that quietly rewrites what your agent is trying to do, and it now runs on a version you cannot pin, diff or roll back. Your eval numbers stop describing a system you control the moment you adopt it.

9 min read

OWASP shipped an interface, not a list

Excessive Agency climbing to third is the headline and the least useful part. The Agent Control Standard is the change: a hook contract a framework fires before a tool call, a memory write or a sub-agent, with an allow/deny/modify verdict behind any policy engine — which turns security advice into something you either implement or do not, and moves the audit boundary into your runtime. It also exposes the number nobody reports: the share of your agent’s effects that pass a hooked call site at all.

11 min read

Discovery is not an inventory

In four days three vendors shipped the same admission: nobody knows what agents are running. CrowdStrike put discovery in the endpoint sensor, AIR raised $50M for an inline firewall at the context boundary, and Tenable and OpenAI put a review in front of a registry. Each answer is complete about one place and silent everywhere else — and every governance regime you are being audited against assumes an authoritative register, not an estimate. The number to start tracking is the gap between the two.

9 min read

An account toggle is not a power of attorney

On 4 September Docusign said its MCP server opens to every agent on 30 September — Claude, ChatGPT, Gemini, Copilot, Slack, any MCP client — governed by account-level admin controls. The law has allowed an automated agent to bind its principal since 1999, on one condition: the act must be attributable to that person. A per-account toggle attributes a class of acts, which is what carried deterministic scripts and is exactly what a model that negotiates strains. Closing that gap is the deployer’s job, and nothing in MCP does it for you.

9 min read

One in ten outages is now AI. That number is not about agents.

The AI share of disclosed outages rose from 1.7% to 10.7% in three years, and agents are not in that denominator — it counts incidents published by AI companies against incidents published by anyone, so it climbs as the sector grows. The figure in the same research that is about agents: 188 of 344 verified enterprise AI incidents had no attacker at all, and the nine documented production deletions share one stage, a credential that outlived the phase it was granted for.

10 min read

MHS vs SiLA 2 vs OPC UA LADS vs ROS 2: the wire format was never the problem

Lab and factory interoperability has been standardised three times already — SiLA 2 since 2019, OPC UA LADS since January 2024, ROS 2 as robotics middleware — and instruments still ship with vendor SDKs, so a fourth spec is not obviously the answer. What Anthropic's Model Hardware Standard adds is the thing none of the three tried: a device that describes its own limits in language a model can read, and a driver that enforces them whichever model is driving. Useful, and not a safety function — keep those apart.

9 min read

n8n vs Dify vs Langflow vs Flowise: the licence names the moat

Flowise archived itself on 13 August 2026 and its maintainers named the reason: coding agents now handle the complexity that a rigid low-code workflow hits a wall on. The three still standing are not surviving on the canvas either — each is defending something underneath it, and each licence says exactly what. n8n forbids offering it to others, Dify forbids multi-tenant operation, Langflow forbids nothing and is owned by IBM. Read the clause before the feature list.

9 min read

Anthropic moved the evidence, not the detector

Enterprise Frontier Safeguards, announced 1 September 2026, resolves a real contradiction: zero data retention forbids the history that cross-session misuse detection requires. Anthropic's fix is to keep the classifier and put the corpus in your own S3, Azure Blob or GCS bucket, under your keys — with alerts routing to you and human review yours by default. That is not only a privacy upgrade. It is a transfer of duty, and the artefact it creates is a discovery-visible record of your own employees' prompts that nobody has written a retention rule for yet.

9 min read

A skipped purchase is not a deployment

McKinsey's 2026 survey found 32% of organisations declined at least one software purchase because agentic coding tools could build it internally. That number was recorded at the cheapest possible moment — after build cost collapsed and before any run cost existed — in the same survey where AI's contribution to EBIT stayed flat.

10 min read

85.5% trust the agent. 41.1% debug it every day.

Temporal surveyed 554 engineers in April and May 2026 and found daily agent use at 80.8%, up from 47.3% a year earlier, with 91.1% reporting improved productivity and 85.5% trusting agent output at least somewhat — alongside 41.1% hitting agent-related issues daily or more and 9.0% continuously. Both sets of numbers are probably accurate, and together they describe a failure rate nobody would accept from a database. The report reads the gap as a state-tracking problem, which is a durable-execution vendor’s reading of a durable-execution question. The more useful reading is that the error handler is a person, and no dashboard has a line for them.

8 min read

Claudeforce runs in two directions, and only one keeps the record inside Salesforce

Salesforce and Anthropic announced one partnership on 26 August 2026 containing two integrations with opposite governance properties. Claude moving into Agentforce keeps the model inside a boundary that already has row-level permissions and an audit log; Salesforce moving into Claude as a plugin moves the session outside it, where the deliberation that produced a write is no longer in the system of record. Both are reasonable products. Buying them as one thing is how a company discovers the difference during its first e-discovery request.

14 min read

Sharing a coding-agent session: every handoff that works throws the transcript away

Claude Code, Codex and Gemini CLI all persist sessions as append-only JSONL, so moving one to another agent looks like a file-conversion problem. It is not. An assistant turn is a claim conditioned on a system prompt, a tool schema, a model and a warm cache that the receiving agent does not have — replay it verbatim and you hand over a false memory. The one converter in the wild strips tool calls into prose on purpose, Anthropic documents its own transcript format as internal and unstable, and Claude Code refuses to resume a hand-copied transcript at all. Four transfer layers, and the useful ones all trade fidelity for something the receiver can re-verify against the repo.

12 min read

The AI AGENT Act Asks for a Record Your Stack Does Not Keep

S. 5051 would require an agent acting for a person to keep real-time records, stay inside its granted authority, and never sub-delegate without explicit permission. Traces record behaviour; all three duties are about permission — which is why the bill hands NIST the job of finding a delegation protocol that does not exist.

9 min read

Microsoft Priced Agent Governance Per Human. Your Fleet Has No Meter.

Agent 365 costs $15 per user per month and nothing per agent, so the one layer of your stack that exists to control fleet growth is also the only layer whose bill ignores it. Two more things do not line up: the licensing unit assumes every agent has a human sponsor, and the inventory can see far more machines than the block button can reach.

9 min read

A2A Moved In With MCP. The Identity Layer Stayed Outside.

On 20 August Google moved A2A into the Agentic AI Foundation, so both protocols in the standard agent stack now share a board, a roadmap and a trademark holder. What they still do not share is a delegation primitive — and the identity work that would supply one is being stewarded at a different foundation entirely.

11 min read

Skill Scanners Read a Different File Than the Agent Runs

Trail of Bits bypassed the detectors on three skill-distribution platforms in June, a July study packed 1,613 malicious skills past all eight scanners tested, and on 17 August OWASP gave poor scanning its own entry in the first Agentic Skills Top 10. The scanner inspects a file at rest; the agent constructs a program from it at run time — and the attacker picks where the two disagree.

10 min read

Slack Code puts the approval in a channel — name one approver anyway

Slack shipped the review surface, not the agent: five partner coding agents you buy separately, working inside a channel with a plan tab, a diff tab and a live preview, and a human approval before anything ships. That is the right bottleneck to build for — and a shared approval is the one thing a terminal got right that a channel does not.

11 min read

MCP Registry vs Smithery vs Docker MCP Catalog vs PulseMCP: four indexes, one missing signal

You can look an MCP server up in four places and get four different kinds of answer: who owns the name, who will host it, who built the image, and what exists at all. Only one of them makes a claim about the artefact you are about to run — and none of them has read the tool descriptions, which is where an MCP server actually attacks you.

9 min read

MCP went stateless, and the state just moved

The 2026-07-28 MCP spec deleted the initialize handshake and the session-id header, so a server can now run behind a plain round-robin load balancer. That operational win is real — but statelessness is a transport property, not a system property. The session did not disappear; its bookkeeping moved onto every request, and the durability that long-running agents actually need came back in through the AWS-contributed Tasks extension as explicit handles. Read the two together before you celebrate a simpler protocol.

9 min read

Stripe bought the meter, not the router

A payments company paid a reported $7 billion for the layer that counts AI usage, seven months after buying the layer that invoices it. Routing was never the scarce asset — the scarce asset is one normalised record of what every model call cost and who it was for, and if that record lives in your request path you are paying a percentage on every step your agents take.

10 min read

August’s Worst Agent CVEs Were Authorization Bugs, and There Was No Patch to Apply

Two agent vulnerabilities scored above 9.0 this month and neither involved a language model. CVE-2026-62830 hit 9.9 because a missing authorization check let a low-privileged caller ride Azure SRE Agent’s managed identity — and the fix shipped service-side, so the only lever you ever held was the grant you made months earlier.

8 min read

GPT-5.6-Cyber Is Gated Because It Refuses Less, Not Because It Knows More

OpenAI's offensive-security model loses to plain GPT-5.6 Sol on both evaluations that score the work product, and wins the one that scores whether it answers at all. Daybreak Red gates a refusal policy, not a capability — which makes patch latency, not model access, the number that should have moved on 10 August.

7 min read

x402 vs AP2 vs ACP vs MPP: The Only Difference That Changes Your Risk

Four agent-payment standards, usually compared on rails. The axis that matters is where the spending cap is stored — a pre-funded wallet, an issuer rule, a one-checkout token, or a mandate the user signed — because that fixes how much a prompt-injected agent can spend before anything else gets a vote.

10 min read

Generative UI Has Two Standards, and They Split Over Who Owns the Catalog

A2UI sends JSON and MCP Apps sends sandboxed HTML — the least consequential difference between them. One has the agent compose components you own, moving the review into your design system; the other installs an interface someone else wrote, moving it to the server boundary. Sort your surfaces by whether you can enumerate them, then pick.

9 min read

Half of Enterprises Scaled Back Their Agents. Seven Percent Can Compute the Ratio.

KPMG found 49% of leaders scaled back an agent deployment over cost and 7% report established ROI — so nine in ten of the organisations that cut did it without a denominator. Cost is metered by a vendor that needs to bill you; value stays at zero until someone builds it. The measurement you cannot add later is the pre-agent baseline.

10 min read

AgentCore vs Foundry vs Vertex AI Agent Engine vs Cloudflare Agents: Nobody Is Selling You the Loop

Two of the four bill the agent loop at about nine cents per vCPU-hour and their prices are 3.6% apart; the other two do not charge for it at all. What each is actually selling is a place to keep the conversation — and AWS closing Bedrock Agents Classic to new customers on 30 July 2026 is the clearest evidence yet about which half of a managed runtime you can afford to rent.

8 min read

Agent Plugins 1.0 Standardises the Bundle and Leaves Trust to Whoever Installs It

Five rival vendors agreed on a directory layout on 6 August, and explicitly declined to agree on install, distribution, permissions, sandboxing or provenance. The format makes one bundle of instructions plus credentialed tool access portable across six clients — which is exactly why the compensating controls are now yours.

11 min read

DeepSeek Is Building a Harness, and the Benchmark Score Already Includes the Scaffold

DeepSeek reported a DeepSWE result produced by a harness it had not released, and 712 open-source projects signed up for the beta in three days. Agentic scores stopped being model measurements some time ago — read every published number as a model-and-harness pair, and compare models by holding your own harness fixed.

11 min read

Google's Agent Calls the Store, and Every Protocol Guarantee Falls Off

Google's shopping agent now phones local shops to check stock — a channel that carries none of the signed identity, scoped authorisation, replay protection or verifiable receipts that AP2 and its rivals were built to provide. The phone is not a stopgap on the way to universal protocol adoption; it is the permanent floor of agent commerce, covering the merchant tail that will never implement an API, and it has no trust primitives at all.

9 min read

Agent Security Just Picked a Layer, and It Is the One You Own

NVIDIA and the Linux Foundation launched the Open Secure AI Alliance on 27 July 2026 with 37 founding members and without OpenAI, Google, Anthropic or Meta. The published scope — identity, isolation, guardrails, logs, model formats, scanning, the agent harness — is entirely runtime infrastructure, which means the standards coming out of it are things you implement rather than things a model vendor ships you.

9 min read

Atlas Shuts Down on 9 August. Agentic Browsing Just Split Into Three.

OpenAI is retiring the ChatGPT Atlas browser nine months after launch and moving its capabilities into a Chrome extension, an in-app browser and a server-side cloud browser. That is not a retreat from agentic browsing — it is the admission that a browser agent never needed a browser. What it needed was proximity to an authenticated session, and the three replacement surfaces are three different answers to whose session it borrows.

8 min read

MCP 2026-07-28: Statelessness Was the Small Part

The 28 July specification retires the initialize handshake and the Mcp-Session-Id header, and every write-up so far has framed that as plumbing. It is not. Dropping the held-open connection forced Sampling, Roots and Logging onto a twelve-month deprecation clock — and those were the features that made an MCP client a peer rather than a caller. The protocol just settled what it is.

10 min read

MCP at 97 Million Downloads: How the Model Context Protocol Won — and What's Still Broken at Scale

Two years from Anthropic's launch, MCP isn't a debate — it's a dependency. Every frontier vendor, every major IDE, and one Pinterest team saving 7,000 engineering hours a month all ship against it. The interesting question is no longer *should you use MCP* — it's what fails at this scale and how the 2026 roadmap plans to fix it.