llm-integration

Claude Inline Tools Beta: What Actually Ships (2026)

September 28, 2026

Claude Inline Tools Beta: What Actually Ships (2026)

Claude's inline-tools-2026-09-15 beta lets you hand the model a brand-new tool mid-conversation — full name, description, and JSON Schema — without touching the tools array or paying to rebuild your prompt cache. It was announced on September 22, 2026, in the same release-notes entry as Claude Opus 5.5.1 A second beta, compact-2026-09-04, documented eight days earlier, lets the API compact a long conversation on demand.

TL;DR: Adding a tool the old way invalidates your cached prefix, because tools sits ahead of everything else in it — a cost one developer measured at 13.5% of their Claude spend across 506 of their own Claude Code sessions.2 The inline tools beta fixes that by moving the tool definition into messages. I read the schema from the docs, then checked it against the actual SDK source rather than the changelog prose, and found the interesting part: the beta header is dated 2026-09-15, but neither official SDK had a type for a by-value tool definition until September 22. That particular gap turned out to cost nothing at runtime in either language — an old client sends the block exactly as written. The neighbouring compaction beta is where an out-of-date SDK actually bites, and only in Python, where an old client rejects the parameter outright with a TypeError rather than forwarding it.

What you'll learn

  • The request shape for adding a tool mid-conversation without invalidating the cache, and the documented rule that keeps the cache intact
  • Why the tool_addition / tool_definition combination is a genuinely different capability from the by-reference tool add/remove announced in July
  • How compact-2026-09-04 compaction and the inline-tools beta hand off tool history through a tool_changes field — and what happens if you forget to combine the beta headers
  • A local mock-server rig that shows exactly what the official SDK puts on the wire, with no API key and no request reaching Anthropic
  • The precise release each capability became typed, measured by grepping SDK source across versions rather than trusting release notes
  • Why an out-of-date SDK costs nothing for an inline tool definition but breaks a Python compaction call outright, measured in both languages

The problem: tools sit ahead of your cache

This is documented behaviour, not a quirk. Anthropic's prompt caching docs state the ordering plainly: "Cache prefixes are created in the following order: tools, system, then messages."3 The same page spells out the consequence in a table of what invalidates what: "Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache" — tools, system and messages levels all together.3

So on a long-running agent that discovers capabilities mid-session, every tool-list change throws away a cached prefix you already paid to write. A Claude Code issue puts numbers on it. The issue author posted a transcript excerpt from one session showing an 857,710-token cache write at 07:47:37, followed at 07:48:49 — about a minute later — by another full 782,873-token rewrite, triggered by tool and MCP-server list changes loading mid-session.2 A commenter on the same issue instrumented 506 of their own local Claude Code session logs (a ~3.2 GB corpus) and reported 2,766 cache-bust events across 191 of those sessions, averaging roughly 81,000 re-billed tokens each, and amounting to 13.5% of the Claude spend in that corpus.2

Two caveats on those figures, since they are the post's only field evidence. They describe Claude Code's own internal tool and MCP listings, and I found nothing confirming that Claude Code uses either beta covered here. And the 506-session analysis is one developer's own machine, self-reported in an issue comment — not a vendor measurement or a fleet-wide audit. What it does establish is that the failure mode is real, recurring, and expensive enough that someone bothered to quantify it.

What's actually new here (and what isn't)

Anthropic shipped a narrower version of this fix two months ago. The release notes entry for July 24, 2026 reads: "Mid-conversation tool changes are now in beta on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5: add or remove tools between turns of a conversation while preserving the prompt cache. Include the mid-conversation-tool-changes-2026-07-01 beta header in your requests."4 So cache-preserving tool changes are not new in September, and neither is the tool_addition block. What July gave you was by reference: point at a tool already declared in tools, or drop one.

September's addition is defining a tool by value — full name, description and schema, inline, for a tool the model has never seen declared anywhere. inline-tools-2026-09-15 covers both, so it supersedes the July header rather than sitting alongside it.5 The shape the docs specify:

{
  "role": "system",
  "content": [
    {
      "type": "tool_addition",
      "tool": {
        "type": "tool_definition",
        "definition": {
          "name": "db_query",
          "description": "Run a read-only SQL query against the analytics database.",
          "input_schema": {
            "type": "object",
            "properties": { "sql": { "type": "string" } },
            "required": ["sql"]
          }
        }
      }
    }
  ]
}

MCP servers can be added the same way, but that is a separate section of the docs with its own requirement: "To add an MCP connector server partway through a conversation, send the mcp-client-2026-09-15 beta header along with inline-tools-2026-09-15."5 Two headers, not one. The response then records each server's fetched tool list in an mcp_tool_listing block, which pins that list when you send it back.1

The caching payoff is the whole point, and the docs are blunt about the mechanism: "The tools array itself never changes, so the cached prefix stays intact."5 Because the definition arrives as an appended message rather than an edit to tools, the earlier prefix stays byte-identical and only the new message is fresh input.

There is one documented exception, and it is easy to walk into: "Keep at least one non-deferred tool in tools. A conversation whose tools array has no non-deferred tool is accepted, but the first tool it defines by value changes the start of the rendered prompt, which costs one full cache miss on that request."5 Declare what you know up front; use inline definitions for what you don't.

A few other limits from the same page, because they fail loudly only after you ship:5

  • Some tool types — the computer-use tool among them — can't be defined in a message during the beta and "return a 400 error that says so"; declare those in tools and add them by reference.
  • cache_control goes either on the content block or inside the definition, never both, and a deferred definition can't carry one at all.
  • Four limits return a 400 with error.details.error_code set to available_tools_limit_exceeded: more than 10,000 deferred tools available after any message; more than 10,000 tools defined after the first user message; definitions sent after the first user message totalling more than 4 MB (4,194,304 bytes); or rendered tool text larger than 4 MB.
  • Tools defined by value are part of the message and are not persisted server-side, so they must stay in the messages array on every later request in the conversation.

When compaction and inline tools meet

The second September beta, compact-2026-09-04, lets you compact a conversation on demand: send a top-level compaction parameter, get back a signed compaction block containing a summary, and replay that block in place of the messages it summarizes.67 Anthropic's changelog documented it on September 14 — ten days after the date stamped into the header's own name — and both official SDKs added typed support the following day, September 15.689

The part that's easy to miss: if a tool was added or removed inside the range of turns you just compacted, that change does not simply vanish into the summary. The docs are explicit, and the behaviour depends on which beta headers were on the compaction request itself, not just on the earlier turns:10

Tool changes inside those turns carry over on their own when the compaction request also carries inline-tools-2026-09-15: the returned block records their net effect in its tool_changes field, so send the block back unmodified. If the block has no tool_changes field, restate those tool changes the same way.

In other words: compact a conversation that added db_query mid-session but forget inline-tools-2026-09-15 on that specific call, and the returned block carries no tool_changes field — you are now responsible for noticing that db_query needs to be back in scope, and for restating it. Include the header and the block carries the net effect forward, to be round-tripped unmodified like its signature.

Two more interactions worth knowing if you combine the betas, both quoted from the same page:10 "A later request can use a different system, different tools, or a different model than the compaction request, and the API still accepts the block. Such a change can invalidate the thinking in the kept turns, but it has no other effect." And: "A system message placed between the block and the kept turns breaks their thinking." So restate standing instructions after the first new user turn instead of wedging them in front of the turns you kept.

Capturing the wire format with no API key

Same methodology as the Claude Opus 5.5 tool_choice post on this site: point the real SDK at a local HTTP server instead of api.anthropic.com, and read what it actually sends. Nothing here reaches Anthropic.

// mock-server.mjs — stands in for the Claude API, and writes each body to disk
// so two SDK versions can be compared byte-for-byte.
import http from "node:http";
import fs from "node:fs";

let n = 0;
const server = http.createServer((req, res) => {
  let body = "";
  req.on("data", (c) => (body += c));
  req.on("end", () => {
    fs.writeFileSync(`captured-${++n}.json`, body);
    console.log("anthropic-beta:", req.headers["anthropic-beta"], "| path:", req.url);

    const parsed = JSON.parse(body || "{}");
    const reply = parsed.compaction
      ? {
          id: "msg_mock_c", type: "message", role: "assistant", model: parsed.model,
          content: [{
            type: "compaction",
            content: "Summary: db_query was added mid-conversation.",
            signature: "mock-signature",
            tool_changes: [{
              type: "tool_addition",
              tool: { type: "tool_definition", definition: { name: "db_query", description: "…",
                input_schema: { type: "object", properties: {} } } },
            }],
          }],
          stop_reason: "compaction", stop_sequence: null,
          usage: { input_tokens: 0, output_tokens: 0,
            iterations: [{ type: "compaction", input_tokens: 144, output_tokens: 276 }] },
        }
      : {
          id: "msg_mock_t", type: "message", role: "assistant", model: parsed.model,
          content: [{ type: "text", text: "(mock reply)" }],
          stop_reason: "end_turn", stop_sequence: null,
          usage: { input_tokens: 1, output_tokens: 1 },
        };

    res.writeHead(200, { "content-type": "application/json" });
    res.end(JSON.stringify(reply));
  });
});
server.listen(8792, () => console.log("capture server on :8792"));

Pointed at that server with @anthropic-ai/sdk 0.128.0 — the current npm version as of this writing — the tool_addition/tool_definition call serializes exactly as the docs describe: the system-role message lands in messages untouched, the pre-existing tools array is unmodified, and the request goes to /v1/messages?beta=true with anthropic-beta: inline-tools-2026-09-15.

Sending both betas on one request — betas: ["compact-2026-09-04", "inline-tools-2026-09-15"] — produces a single header reading compact-2026-09-04,inline-tools-2026-09-15: comma-joined, no space, which is the behaviour a bug fix in the immediately preceding release (0.127.0, "join multiple anthropic-beta values with a comma and no space") was written to guarantee.8

One negative result worth recording, because it corrects the obvious way to test this. The response parser does expose the mocked tool_changes field — result.content[0].tool_changes[0].tool.type reads back as tool_definition. But that proves nothing about typed support: I ran the same probe through @anthropic-ai/sdk 0.105.0, which has no tool_changes type anywhere in it, and read back the same nested value. These clients pass unknown response fields straight through, so a runtime property check cannot distinguish "the SDK models this field" from "the SDK ignores it." The only evidence that actually settles typed support is the source itself, which is the next section.

When could an SDK actually build this for you?

"Added support" in a release note can mean anything from a fully generated type to accepting a string. So I downloaded each relevant release of both official SDKs, extracted it, and grepped for the generated type names — with a positive control in every run (a token that must be present, like BetaMessage) so an empty result couldn't be mistaken for absence.

The result is more granular than either changelog suggests. Compaction types are not new at all: the oldest release I checked, TypeScript 0.105.0 from 2026-06-18, already carried an unsigned BetaCompactionBlock and an autocompact config with pause_after_compaction. What it did not have was any way to ask for a compaction — no top-level compaction parameter, no {type:'summarize'} config. Each piece arrived separately:

DateTypeScriptPythonWhat became typed
earliest checked0.105.0 (06-18)1.4.0 (09-04)Unsigned BetaCompactionBlock + autocompact config — but no on-demand compaction parameter
2026-07-240.115.0—By-reference tool_addition/tool_removal blocks, the day the July beta was announced4
2026-09-14——compact-2026-09-04 documented publicly6
2026-09-150.126.01.6.0The top-level compaction?: BetaCompactionConfig parameter, signature on the compaction block, and the compact-2026-09-04 header string
2026-09-180.127.01.7.0No new surface here; TypeScript got the beta-header comma-join fix8
2026-09-220.128.01.8.0By-value tool_definition, the inline-tools-2026-09-15 header string, and tool_changes on the compaction block

That last row is the one that matters if you are building the handshake described above. tool_changes — the field that carries tool history across a compaction boundary — became typed on exactly the same release as the by-value tool definition it refers to, in both languages, on September 22. Before that, the compaction block type had a signature but no tool_changes at all.

Both changelogs describe the September 15 work in identical words — "add compaction parameter and signed compaction blocks (beta)" — which tracks, given both SDKs are generated rather than hand-written (each emits X-Stainless-* request headers).89 Scope note on the archaeology: I checked TypeScript 0.105.0, 0.115.0, 0.120.0, 0.124.0, 0.125.0, 0.126.0, 0.127.0 and 0.128.0, and Python 1.4.0 through 1.8.0 (every release in that range). tool_definition and the inline-tools-2026-09-15 string are absent from every one of those before September 22 and present in both SDKs from that date. I have not checked the seven TypeScript releases between 0.116.0 and 0.123.0 other than 0.120.0.

The old-SDK gap was not the same in both languages

None of that means the feature didn't work before September 22, and this is where the two ecosystems diverge in a way no release note mentions. I pointed @anthropic-ai/sdk 0.105.0 — which contains no reference to tool_addition, inline-tools or tool_definition anywhere in its types — at the same mock server and sent the identical request object the current SDK had just sent. Both bodies arrived at 717 bytes and compared byte-for-byte identical. The old client also passed a top-level compaction: {type: "summarize"} through untouched, despite having no type for it.

Python behaves differently, because its create methods take explicit keyword arguments rather than an options object. Real output from the two probes against anthropic 1.5.0:

Python SDK version: 1.5.0

--- TEST 1: inline tool_addition / tool_definition on OLD Python SDK ---
RESULT: SUCCEEDED at runtime, id = msg_mock_t

--- TEST 2: top-level compaction={'type':'summarize'} on OLD Python SDK ---
RESULT: TypeError -> Messages.create() got an unexpected keyword argument 'compaction'

Re-running test 2 on anthropic 1.6.0 — the release that added the parameter — succeeds, which fixes the boundary at that version rather than leaving it inferred from the changelog.

So the practical rule is per-language and per-location. Message content is a plain dict or object that both SDKs, in the versions I tested, serialize as written — which is why a hand-built tool_addition with a full tool_definition worked on both old clients. A new top-level parameter is different: JavaScript spreads it into the request body regardless, while Python rejects it at the call site until the version that adds the keyword. A Python service wanting on-demand compaction needed anthropic ≥ 1.6.0 to make the call at all — not merely for editor support.

Practical checklist

  • Keep at least one non-deferred tool declared up front. If tools would otherwise start empty, the first by-value addition costs one full cache miss anyway.5
  • Adding an MCP server mid-conversation needs two headers, mcp-client-2026-09-15 alongside inline-tools-2026-09-15, and the mcp_tool_listing block must be sent back to pin the list.15
  • Put cache_control in one place only — the block or the definition, never both — and remember a deferred definition can't carry one.5
  • If you compact a conversation that changed tools mid-range, include inline-tools-2026-09-15 on the compaction request itself, not just on the earlier turns, or the returned block won't carry a tool_changes field at all.10
  • Restate standing system instructions after the first new user turn following a compaction, never in a system message between the compaction block and the turns you kept.10
  • Pin by what you actually need. anthropic (Python) ≥ 1.6.0 is a hard runtime requirement to pass compaction at all; ≥ 1.8.0 and @anthropic-ai/sdk ≥ 0.128.0 are what give you real types for a by-value tool_definition and for tool_changes.

What this does not prove

No API key was used anywhere in this post and no request reached Anthropic. The request and response shapes, limits and cache rules are quoted from the official docs, fetched 2026-09-28;1356710 what I measured is what these SDK versions put on the wire against a server I control, and which types exist in which published release. I have not observed a live API response, so I have not confirmed that the real API returns a tool_changes field in practice, only that the docs specify it and the current SDKs type it — my tool_changes payload was a fixture I wrote myself.

The cache-cost figures are one developer's self-reported measurements of Claude Code, not the Messages API, and not independently reproduced. The version archaeology covers the two official SDKs at the versions listed above; I did not check LangChain, the Vercel AI SDK, or any gateway for inline-tools or compaction support, and I did not test whether those forward or reject an unknown top-level parameter the way these two clients do.

Bottom line

The date in a beta header tells you when Anthropic cut the API change, not when you can build against it, and not what happens if you don't. compact-2026-09-04 was documented ten days after its own date stamp and typed the day after that. inline-tools-2026-09-15 waited a week for its announcement and its first typed helper, which arrived together on September 22 with Claude Opus 5.5.

The more useful finding is what an out-of-date SDK actually costs you, because it is not one answer. Inside messages, nothing: both languages' old clients sent a by-value tool definition byte-for-byte correctly, types or no types. At the top level, everything: the same Python client that happily forwarded a tool_addition block refused to accept a compaction argument at all. If you are adopting either beta, check which of those two shapes your change is before deciding whether your lockfile matters.

For what an MCP-tool-list mismatch looks like from the client's side of an active session, see AI SDK tool drift detection against a real MCP server. And for the cross-session side of context management — memory that survives past a single conversation rather than a single compaction — see Claude Memory Tool + Context Editing.

Footnotes

  1. Anthropic, Claude Platform release notes — the September 22, 2026 entry announcing Claude Opus 5.5 and inline tool definitions, including the mcp_tool_listing pinning behaviour. https://platform.claude.com/docs/en/release-notes/overview (fetched 2026-09-28) ↩ ↩2 ↩3 ↩4 ↩5

  2. anthropics/claude-code issue #92033, "Mid-conversation tool/MCP list changes silently invalidate prompt cache, causing repeated full-price rewrites within minutes" — the 857,710/782,873-token rewrite pair is from the issue author's own transcript excerpt; the 506-session, ~3.2 GB corpus analysis (2,766 events across 191 sessions, ~81,000 tokens each, 13.5% of spend in that corpus) is from commenter u2giants' own local logs. The issue's open/closed state was not legible in the page as fetched, so no status is claimed here. https://github.com/anthropics/claude-code/issues/92033 (fetched 2026-09-28) ↩ ↩2 ↩3

  3. Anthropic, "Prompt caching" — the tools → system → messages prefix order and the cache-invalidation table quoted here. https://platform.claude.com/docs/en/build-with-claude/prompt-caching (fetched 2026-09-28) ↩ ↩2 ↩3

  4. Anthropic, Claude Platform release notes — the July 24, 2026 entry announcing mid-conversation-tool-changes-2026-07-01, quoted in full. https://platform.claude.com/docs/en/release-notes/overview (fetched 2026-09-28) ↩ ↩2 ↩3

  5. Anthropic, "Mid-conversation system messages and tool changes" — the tool_addition/tool_definition shape, the cache-preservation rule and its non-deferred-tool exception, the two-header MCP requirement, the computer-use restriction, and the available_tools_limit_exceeded limits. https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages (fetched 2026-09-28) ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12

  6. Anthropic, Claude Platform release notes — the September 14, 2026 entry documenting on-demand compaction under compact-2026-09-04. https://platform.claude.com/docs/en/release-notes/overview (fetched 2026-09-28) ↩ ↩2 ↩3 ↩4

  7. Anthropic, "Compaction on demand" — the compaction request parameter, the signed compaction block, and per-iteration usage billing. https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand (fetched 2026-09-28) ↩ ↩2

  8. anthropics/anthropic-sdk-typescript CHANGELOG and tagged releases 0.115.0 through 0.128.0; version timestamps from the npm registry (npm view @anthropic-ai/sdk time). Type presence verified by extracting each published tarball. https://github.com/anthropics/anthropic-sdk-typescript/blob/main/CHANGELOG.md (fetched 2026-09-28) ↩ ↩2 ↩3 ↩4 ↩5

  9. anthropic (Python) release history on PyPI, versions 1.4.0 through 1.8.0, cross-checked against the anthropics/anthropic-sdk-python v1.6.0 release notes. Type presence verified by extracting each published wheel. https://pypi.org/pypi/anthropic/json, https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.6.0 (fetched 2026-09-28) ↩ ↩2 ↩3

  10. Anthropic, "Compaction: thinking blocks and tool changes" — the tool_changes field and its dependency on inline-tools-2026-09-15, plus the system/tools/model swap and system-message placement rules, all quoted verbatim. https://platform.claude.com/docs/en/build-with-claude/compaction-thinking-blocks (fetched 2026-09-28) ↩ ↩2 ↩3 ↩4 ↩5 ↩6

Frequently Asked Questions

It lets you send a role: "system" message mid-conversation containing a tool_addition block whose tool is a full tool_definition — name, description and JSON Schema — for a tool the model has never seen declared. Because the tools array itself never changes, the cached prefix stays intact. 5