AI Blog

Safari MCP vs Chrome DevTools MCP vs Playwright MCP vs extension agents

Tool counts decide nothing here. The axis that determines both whether a browser agent can do the job and how bad a hostile page gets is which session it holds — an isolated automation context, a dedicated profile quietly accumulating logins, or your own signed-in browser. Both browser vendors that shipped an MCP server this year deliberately kept your own session out of it, which is why neither does the agentic-shopping demo everyone expected.

By Agentic AI Wiki 13 min read

Every comparison of browser tooling for agents leads with tool counts, and the tool count decides nothing. The axis that determines whether the thing works at all, and how bad it is when a page turns hostile, is which session the agent gets: no identity, a dedicated profile it accumulates itself, or the logged-in browser you use for your bank. Apple, Google, Microsoft and the extension vendors each picked a different answer, and once you sort them that way the choice stops being about features and starts being about what you are willing to hand a model that just read an untrusted page.

At a glance

Four ways to give an agent a browser, ordered by how much of you it holds.

OptionWhat it isSession the agent getsStrongest at
Safari MCP Apple's own server, safaridriver --mcp, stable since Safari 27 An isolated automation session — no cookies, passwords or history Debugging WebKit rendering on the browser you cannot emulate
Chrome DevTools MCP Google's wrapper over the DevTools Protocol A dedicated Chrome profile, persistent across runs Performance traces and network-level diagnosis
Playwright MCP Microsoft's server, acting on the accessibility tree A dedicated profile by default; --isolated for clean rooms Driving flows reproducibly, across three engines
Extension agents An in-browser assistant, e.g. Claude in Chrome Your real profile, with every login you already have Tasks on sites you are signed into and cannot script
How much of your identity the agent holds Horizontal bars for four options. Safari MCP is shortest, running an isolated automation session with no cookies or logins. Chrome DevTools MCP and Playwright MCP are mid-length, each using a dedicated profile that accumulates logins across runs. Extension agents are longest, holding the user's own profile and every existing session. Share of your identity the agent holds Safari MCP isolated session — no cookies, no logins Chrome DevTools MCP dedicated profile, reused across runs Playwright MCP dedicated profile, or --isolated Extension agents your own profile — every session you have less capable, fails safe more capable, fails open
The same ladder measures capability and exposure. There is no rung that gives you one without the other.

Why the session is the whole decision

Three session models and what an injected instruction reaches Three columns: a clean automation context holding no credentials, a dedicated profile that accumulates logins over time, and the user's own profile carrying every existing session. Each column lists what the context holds and what an instruction injected into a visited page can therefore reach, with a closing note that the reach is the cookie jar. One injected line, three cookie jars Clean automation context Dedicated profile Your own profile HOLDS no cookies, no saved passwords, no history, no AutoFill whatever you have logged it into, accumulating, no lifecycle every session you have — mail, bank, repos, admin consoles CAN DO THE ERRAND? No — it is not you anywhere On the systems you signed it into Yes, anywhere you are signed in Page carries an injected instruction SO THE INJECTION REACHES public pages, the local filesystem if exposed, and nothing of yours everything that profile has ever been logged into, reviewed by nobody account takeover as you, acting as you Capability and exposure are the same variable; the cookie jar is the dial.
An injected instruction reaches exactly as far as the cookie jar the agent is holding.

Capability and exposure are the same variable

An agent in a clean context can load a public page, read the DOM, run JavaScript and screenshot it. It cannot check your order status, because it is not you. Give it your profile and it can do the whole errand — and a paragraph on any page it visits now addresses a process holding every session you have. Prompt injection stops being a scraping nuisance and becomes an account-takeover primitive, not because the model got weaker but because the blast radius grew.

This is why both browser vendors that shipped an MCP server this year kept your own session out of it. Apple went to the bottom of the ladder outright: a dedicated automation session isolated from your browsing, with no access to cookies, saved passwords, AutoFill or history. Google stopped one rung up, launching a separate Chrome profile under its own cache directory rather than attaching to yours. Neither is an oversight or a v1 limitation — it is the only defensible default for a server that any MCP client can attach to, and it is also the reason neither one does the agentic-shopping demo people expected from "the browser now speaks MCP."

The middle rung is the one people misread

A dedicated profile is not a clean room. Chrome DevTools MCP reuses its user-data directory between runs unless you ask for isolation, and Playwright MCP is persistent by default too. That is deliberate and convenient: you log into your staging environment once and the agent stays logged in. But it means the profile accumulates credentials over weeks, nobody reviews what is in it, and it is on disk in a cache directory with no lifecycle. Six months in, "a dedicated profile" can hold more sessions than anyone remembers granting.

Safari MCP — the only way to see WebKit, and deliberately powerless

What it does

Apple ships the server as a mode of safaridriver, first in Safari Technology Preview and now in stable Safari 27. It exposes on the order of seventeen tools covering the debugging loop: screenshots, page content, JavaScript evaluation, click/type/scroll/hover interactions, buffered console messages, network request listing and detail, an accessibility audit, and media-feature emulation for things like dark mode and reduced motion. It runs as a local stdio subprocess and makes no network calls of its own.

Getting it on is two toggles deep

You enable web-developer features in Safari's Advanced settings, then allow remote automation and external agents in the Developer pane. That friction is the enterprise story too: it is a per-user setting in a consumer browser, which is a thin basis for a fleet policy, and worth knowing before you plan around it.

Who it fits

Anyone debugging a real WebKit bug. Playwright's WebKit build is close, but it is not Safari on macOS or iOS, and the class of defect where that distinction matters is exactly the class you cannot reproduce anywhere else. Outside that, the isolated session and the smaller tool surface make it the least useful of the four.

Chrome DevTools MCP — the diagnostic one

What it does

Google's server is a thin, standard wrapper over the Chrome DevTools Protocol, driving Chrome through Puppeteer, with a tool surface several times larger than Apple's — roughly sixty tools. The differentiated half is not the clicking; it is performance traces with extracted insights, network request inspection and console access. If the question is "why is this page slow, or noisy, or failing a request", this is the one with direct leverage, because it hands the agent the same data a human opens DevTools for.

The profile detail that matters

By default it starts stable Chrome against its own user-data directory under the user's cache path, reused across runs, with only one browser allowed to use it at a time. Set the isolated option and you get a temporary directory that is cleared when the browser closes. For CI you want isolated; for a long-running debugging session you usually do not, and that is the decision to make explicitly rather than inherit.

Who it fits

Front-end and performance work, Chrome-only, where the agent's job is to explain a page rather than to complete a flow on it. It is also the weakest of the four at reproducibility, because a persistent profile plus a real browser is a lot of hidden state.

Playwright MCP — the one that drives, and the one that is honest about state

What it does

Microsoft's server operates on the page's accessibility tree rather than on screenshots, which is a bigger deal for agents than it sounds: the model receives structured text with roles and names, so it selects elements semantically instead of guessing pixel coordinates, and every observation is cheap, diffable and loggable. It runs Chromium, Firefox and WebKit, which makes it the only cross-engine option here.

State, explicitly

Profiles are persistent by default, stored per channel and workspace under the platform cache directory so different projects get different profiles automatically, and overridable with --user-data-dir. Pass --isolated and each session starts fresh and discards its storage on close; pair it with --storage-state to seed exactly the logged-in state a test needs and nothing else. There is also an extension mode that attaches to a tab in your own browser, which moves it up the ladder to the top rung with all that implies.

That triple — persistent, isolated, or seeded from a file — is the most complete answer to the session question any of these four gives, and seeded isolation in particular is the shape you want for evals: reproducible, credential-scoped, and re-creatable from a file you can review.

Who it fits

Anything where the agent is meant to accomplish a flow rather than describe a page: end-to-end testing, QA on release, form-driven workflows, and any agent that must run on more than Chrome. It is the default choice, and the accessibility-tree approach is also why its traces are the easiest to keep — see screenshots and DOM artefacts for why that matters more than it looks.

Extension agents — the only ones that can do the errand

What they do

An in-browser assistant such as Claude in Chrome, generally available on paid plans since late August 2026, reads the page you are on and acts on it — clicking, typing, navigating, filling forms — using the logins you already have. That last clause is the entire product. It is also the reason nothing on the lower rungs can substitute for it: no automation profile knows your account.

The security record is the argument, not a footnote

Two disclosed flaw classes in the Claude extension make the exposure concrete rather than theoretical. ShadowPrompt chained an overly permissive origin allowlist — any subdomain matching *.claude.ai could submit a prompt for execution — into zero-click injection from a visited page; Anthropic tightened it to an exact-origin check in extension version 1.0.41 after disclosure in late December 2025. Separately, ClaudeBleed turned on the observation that any installed Chrome extension could send commands, because trust was placed in the origin of a command rather than in its execution context.

Both were fixed. The structural point survives the fixes: an agent operating inside your trusted session executes as you, so a successful injection reads files, mail and code, acts on your accounts, and can tidy the interface afterwards. Everything on the lower rungs fails safe by not having the credentials; this rung fails open by design. Use it for tasks you would have done yourself in that tab, on sites you chose, and do not leave it attached while browsing the open web.

The axes that actually separate them

The three axes that separate browser tooling for agents Three axes; tool count is not one of them Which session What it observes Reproducible state Decides: how far an injected line on any page can reach Decides: reliability across a redesign, and trace size Decides: whether a failure can be re-run and scored Check: open the profile and list what it is logged into Check: is a step an a11y tree, a screenshot, or a CDP trace? Check: can a colleague re-run this failure from a file? Isolated → dedicated → yours Pixels → tree → protocol data Persistent → isolated → seeded Answer these three and the shortlist is one; the tool list never separated them.
Session decides exposure, observation surface decides reliability, and state decides whether a failure can be re-run.

Beyond the session, the second real axis is what the agent observes. Pixels are universal and expensive: a vision model, a large artefact per step, and coordinates that break on a layout change. An accessibility tree is text, so it is cheap, semantic and greppable, and it is why Playwright's selections survive a redesign more often. Protocol data — traces, requests, console — is neither, and it is the thing only the DevTools server really gives you.

The third is reproducibility, which follows directly from the profile decision and is where teams get burned. A persistent profile makes a failure you cannot re-create on another machine; seeded isolation makes one you can. For any run you intend to score, that is not a preference.

When to pick which

SituationPickWhyWatch for
Agent must complete a flow, possibly on Firefox or WebKitPlaywright MCPAccessibility-tree actions, three engines, explicit state modesPersistent profile as the default
"Why is this page slow / erroring"Chrome DevTools MCPPerformance traces and network data no other option hasChrome-only; hidden profile state
A genuine WebKit or Safari-specific bugSafari MCPThe only way to inspect the real engineIsolated session; per-user toggles
An errand on a site you are logged intoExtension agentNothing else has your sessionInjection is account takeover; scope the sites
Evals or CI you intend to scorePlaywright MCP, isolated + seeded stateReproducible and credential-scopedStorage-state files are secrets

FAQ

Does Safari shipping an MCP server mean agents can now shop on my behalf in Safari?

No, and that is by design. The server drives a dedicated automation session with no access to your cookies, saved passwords, AutoFill or history, so it cannot act as you on any site. It is a developer debugging tool that happens to speak a protocol consumer agents also speak.

Is a "dedicated profile" safe to treat as a sandbox?

Not by itself. It is a separate cookie jar, which is real isolation from your personal identity, but it persists between runs and accumulates whatever you log it into. Review it like a shared service account: know what is in it, keep production credentials out of it, and use isolated mode for anything you want reproducible.

Accessibility tree or screenshots — which should an agent act on?

The tree, wherever it exists. It is cheaper, semantic, survives visual redesigns better, and leaves a trace you can redact and search. Screenshots are the fallback for canvas-heavy apps and for confirming what the page actually looked like at a decision point.

Can I run more than one of these at once?

Yes, and it is a common setup: Playwright MCP to drive, DevTools MCP to diagnose. Watch the profile locks — the DevTools server allows only one browser per user-data directory — and watch your context budget, since two servers can add well over a hundred tool definitions to every request.

What is the one control to add before letting any of these near real sites?

Egress and destination limits, not a better model. Decide which origins the browser may reach, treat every page as untrusted input, and require confirmation for irreversible actions — because the page is the attack surface regardless of which of the four you chose.

Further reading

On this wiki:

Project documentation: