AI news ·
Arize ships hosted MCP server for AX alongside CLI and skills
Arize shipped a hosted MCP server exposing 45 tools to agents in Claude Code, Cursor, and Claude Desktop without leaving the editor. Agents using structured tool manifests made 38% fewer invocation errors than those composing free-form shell commands.

Arize has shipped a hosted MCP server for its AX observability platform, giving agents in Claude Code, Cursor, and Claude Desktop direct access to 45 tools without leaving the editor. The move comes less than a year after the company's CPO called MCP a "context tax" - a criticism that was accurate at the time but has since been resolved by improvements in agent harnesses.
The new server lets an agent pull the ten slowest traces for a project, diff two experiment runs to find regressions, read evaluator prompts, check active monitors, and look up dataset examples. There is nothing to install. Users point their client at a regional endpoint and pass an API key as a bearer token.
Why the context tax argument no longer holds
For most of the past year, MCP tool definitions loaded into context at the start of every agent session - every tool, every parameter, every description - before the model read a single message. Community measurements put a Claude Code session with five to ten servers installed at 50,000 to 67,000 tokens of overhead. A GitHub plus Slack plus Sentry plus Grafana plus Splunk setup landed around 55,000 tokens in tool definitions alone. Every connected server paid a tax whether it was used or not.
CLIs never had that problem. A command string costs a few dozen tokens, agents already know git, docker, and curl from pre-training, and output composes in the shell through pipes and grep. But the harnesses caught up. Claude Code, Cursor, and Codex now do tool search: only server names sit in context at startup. Full tool schemas load when the model goes looking for one. Anthropic's engineering team cut token usage from roughly 150,000 to 2,000 on a Drive-to-Salesforce task by having the agent write code against MCP servers instead of calling tools one at a time.
Anthropic's Thariq Shihipar said he was not expecting it, but "MCPs are now better than CLIs for most integrations." Models have improved at tool calling, hosts defer tool loading, and MCP is stateless over HTTP. Composition shifted too: instead of piping output through grep, agents put a filter parameter on the tool and compose inside one call.
Three surfaces, one API
Arize built three interfaces on the same REST API. The MCP server exposes 45 named tools plus server_info across projects, traces, datasets, experiments, evaluators, monitors, and prompts. The ax CLI ships 22 command groups covering the full platform surface. And a set of skills, installed with ax skills install, teaches an agent how to do a job with the CLI.
A skill is a SKILL.md file plus optional scripts and references. It costs roughly 100 tokens of context until invoked - only the name and description load at startup. The full instructions, including gotchas and procedural knowledge no tool schema carries, arrive only when the task matches. The arize-experiment skill encodes the actual workflow: verify the environment, resolve the space and project, list experiments with the right flags, then compare runs.
The MCP project finalized a Skills extension in September that lets a server publish skills as resources. A client calls skills/list, gets back names and descriptions, and reads the full SKILL.md only when the model picks one. Host support is still rolling out, so Arize's skills currently install through the CLI. The company expects the AX MCP server to serve skills directly once clients catch up.
A decision framework for teams
The choice between MCP, CLI, and skills depends on where the agent runs and who is asking. A PM in Claude Desktop asking which experiments regressed on hallucination has no terminal. A support engineer in Cursor wants to pull failed traces without leaving the editor. Claude Desktop, Cursor, and Zed all speak MCP natively - for those surfaces, MCP is the only option that does not require dropping into a shell.
CLIs handle scripted pipelines, CI jobs, cron tasks, Lambda functions, and Docker entrypoints where no interactive session exists. They are also the foundation skills are built on. Most production agents use both: CLIs handle execution-heavy, scriptable work, while MCP supports integration and discovery for clients without shell access.
Structured tool manifests reduce errors. One comparison found agents using them made 38% fewer invocation errors than agents composing free-form shell commands. An MCP server also exposes a fixed set of named operations - a host can grant or deny each one, and the model never holds raw credentials in its context.
Why this matters for operations and development teams
For IT, operations, and development professionals, the practical takeaway is that MCP is no longer a context-cost liability. Modern harnesses load tool schemas on demand, and the decision now hinges on access patterns: MCP for any agent running without a shell, CLI for scripted pipelines and non-interactive environments, and skills to encode the procedural knowledge that turns API access into competent task execution. The three surfaces share one API, so behavior stays consistent whether an agent calls AX from a chat client, a shell script, or a CI pipeline. Teams building internal agent tooling should expose named, annotated tools with filters rather than hiding everything behind generic search-and-execute pairs - clients already handle tool discovery, and a custom layer per server forces the model to navigate two discovery paths.