17°
Portada del artículo: AWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloud
AgentsLangGraphAWSBenchmarksLLM

AWS says its agent harness saves 28% of tokens. I measured 1.82x more expensive locally, 88% cheaper in the cloud

I tested Strands Harness, AWS's agent runtime, against a bare LangGraph agent: same 3 tools, same gatekeeper, same 20-turn conversation. Against a local Qwen3.8-27B with a 16k context window it spent 1.82 times more tokens. Against the same job on a cloud model, with temperature finally matched between the two, it spent 88.8% less. And three rounds of adversarial audit had to correct my own mistakes along the way: 11 extra tools, a turn limit that lied, and a cache double-count that inflated the first number 4x.

Efrain Garay 23 September 2026

Playing summary

I wanted to test whether AWS’s new context management, Strands Harness, solves something real. I had the setup on hand by chance: I had already broken and rebuilt a chat on Amazon Bedrock AgentCore, with LangGraph, a deterministic gatekeeper, and a 64-turn adversarial conversation already recorded. The plan looked simple: swap only the harness, keep everything else the same, measure tokens.

It was not simple. I ended up auditing my own experiment three times, and the first two conclusions I was about to publish were both false.

In 57 seconds: the 28% AWS claims, the 1.82x more expensive I measured locally, the flip to 88.8% cheaper in the cloud with temperature matched, and the 918,648 tokens Claude Code burned as a third harness.Watch it in the reel viewer →

What Strands Harness promises, and what it never touches

Strands Harness is not a new AgentCore feature: it is a separate SDK (pip install strands-harness), announced on September 21 with a specific number: 28% lower token consumption than other harnesses, achieved by truncating tool outputs past 1,500 tokens and auto-summarizing at 85% of the context window.

My first idea was to test it against the jailbreak that broke my AgentCore chat: if context management is better, does it survive a long adversarial conversation better? I checked the code before wasting time: that incident’s real fix lives in scope.py, a regex classifier that runs before the message ever touches the model. Of the 64 turns in that conversation, 9 actually reach SITE and invoke the LLM; the other 55 get a fixed reply without spending a single token. What stops the jailbreak is the classifier, not the context window: any harness’s context management is irrelevant to that.

I changed the question: I took a real 20-turn conversation (not adversarial, from the same eval suite as the chat: someone who says hi, asks scattered things, wants a summary, claims “the site only has 2 articles” when it has 41), the same system_prompt, the same 3 tools (list_posts, search_posts, read_post) and the same gatekeeper. I ran it against two backends: a Qwen3.8-27B served locally with llama.cpp (16,384 tokens of real window) and MiniMax-M3 in the cloud, the model the production chat actually uses.

The matrix: two harnesses, two backends, the same conversationSame system prompt, same 3 tools, same gatekeeper, the same 20 turns.
real conversation · 20 turns3 tools · 1 gatekeeperLangGraphcreate_react_agentStrands Harnesscontext_manager=autoQwen3.8-27B local · 16k ctx181,833 tokensMiniMax-M3 · cloud103,952 tokensQwen3.8-27B local · 16k ctx331,118 tokens1.82xMiniMax-M3 · cloud11,618 tokenstemperature matched

Against the local model, Strands spent 1.82 times what a bare ReAct agent did. Against the cloud model, once temperature was finally matched, it spent a fraction of what LangGraph did — the Anthropic cache breakpoint Strands inserts and LangGraph does not.

First run: both crash at the same point. It was a lie.

The first time, LangGraph and Strands appeared to die at the same point against the local model, with the same kind of window-exceeded error. I was about to publish that. I asked two adversarial reviewers (Codex, and a second model separately) to audit the experiment before writing another word, because that same week I had already found invented numbers in another post on this site and did not want to repeat it with my own data.

They found three mistakes of mine, not of Strands:

  1. create_harness() adds 11 extra tools by default (shell, files, background tasks, subagents), on top of the 3 I did ask for: 14 total per request. builtin_tools=[] only turns off the coding-agent ones; the system plugins and the harness’s own contract stay on top regardless. That inflates every request without it being obvious.
  2. An artificial limits={"turns": 8} I had set to stop a runaway loop silently returned stop_reason="limit_turns". My script counted those turns as completed when they had actually been cut short.
  3. LangGraph, in the original run, actually completed all 20 turns without crashing. The error I saw first came from an earlier, different run that I had confused with this one.

I fixed all three: dropped the harness contract (system_prompt=SYSTEM_PROMPT instead of instructions, which skips the default preamble), turned off the system plugins (builtin_plugins=[]), removed the turn limit, and added real truncation detection (finish_reason/stop_reason) on both sides, not just one.

Second run: 4.1 times more expensive. Also a lie, halfway.

With those fixes, Strands completed all 20 turns against the local model without crashing. Good news. But the total, 752,659 tokens against LangGraph’s 181,833, came out to 4.14 times more expensive. I sent this out for audit again, with the new numbers.

They found a fourth bug, subtler this time: Strands’ llama.cpp provider reports inputTokens as the full prompt, and separately exposes cacheReadInputTokens as an informational subset of that same total, not something additional. My script was adding both. Strands’ own SDK already ships the correct function for this (_total_prompt_tokens, which decides per provider whether cache goes inside or on top), and I had reinvented it wrong.

With the accounting corrected (the SDK’s real function, not mine), the number dropped to 331,118 tokens: 1.82 times more than LangGraph, not 4.1. Still a real, large difference, just less than half of what I was about to publish.

Total tokens for the identical 20-turn conversationtotal tokens · lower is better
  1. LangGraph · local Qwen181,833
  2. Strands · local Qwen331,118
  3. LangGraph · MiniMax cloud103,952
  4. Strands · MiniMax cloud11,618

Local Qwen3.8-27B Q3_K_M (llama.cpp, 16,384-token window) and MiniMax-M3 cloud. Same system prompt, same 3 tools, same gatekeeper, same 20-turn conversation. Final accounting after 3 rounds of adversarial audit.

Against the local model: the 16k window explains a good part of it

Neither harness crashed this time, but both ended up close to the limit. LangGraph accumulated up to 15,314 tokens of context by turn 19 (93% of the window), with zero context management at all: simply because my baseline discards tool results between turns and only keeps human/assistant text, a much more aggressive memory policy than Strands, which retains and compacts.

Strands, with context_manager="auto" active, reached 16,370 of 16,384 (99.9%) by turn 18, and its own auto-summarizer never fired in time: the threshold is 85% utilization, but that calculation only counts messages, not the system prompt or the tool specs, so it underestimates real usage.

How long each harness takes to spend its total, turn by turnAgainst the same local Qwen3.8-27B. All 20 turns, scaled to reading speed.
LangGraph181.833 s
Strands Harness331.118 s

In thousands of accumulated tokens. Strands finishes with 1.82 times more, even without hitting the window.

The same conversation, the same model, the same 16,384-token window. The difference is not that one crashes and the other does not: it is how much each one spends to get to the same place.

Part of that gap has a concrete, checkable cause: counting the tools the model actually receives per request (with a direct probe against the server), Strands sends 6 tool specs per call, not 3: on top of list_posts/search_posts/read_post, the harness’s own ContextOffloader adds retrieve_offloaded_content, the SDK’s context manager adds retrieve_context, and background-task handling adds a third. Measured on the real server: 6 tools cost 2,549 header tokens per request; 3 cost 1,324. Multiplied across the 38 internal cycles Strands ran over the whole conversation, that alone explains a meaningful chunk of the 1.82x. I could not strip those three tools out without dismantling the very mechanism I was testing: they are exactly the context-management machinery.

Against the cloud model: the result flips

Here is the twist I almost failed to measure correctly. The first run against MiniMax gave Strands a 12% higher cost than LangGraph, close to a tie. But that run carried an asymmetry I had not closed: LangGraph sends temperature=0.2 explicitly; Strands’ AnthropicModel, when given the same parameter, made the installed Anthropic SDK fail (AsyncMessages.stream() got an unexpected keyword argument 'temperature', a version mismatch, not a Strands bug). I left it without that parameter and moved on, flagging it as an open gap.

A second reviewer found the way through: the Anthropic SDK does accept extra_body, and Strands forwards params wholesale. With params={"extra_body": {"temperature": 0.2}}, the call works the same as LangGraph’s.

With temperature finally matched, the result did not just get closer: it flipped entirely. Strands went from 116,472 down to 11,618 tokens, 88.8% less than LangGraph’s 103,952.

Strands Harness announcement · Sep 21 2026 · against my own bench
1/ 1se sostienen al medirlas
Token reduction vs. another harnessse sostiene
dicen 28% less

88,8% lessMeasured against LangGraph with the identical 20-turn conversation, MiniMax-M3, temperature matched.

AWS compares against Claude Code and Codex, not against LangGraph, so this is not the exact same comparison. Even so, the number measured on my own bench far exceeds what was published — in the opposite direction from what I first measured against the local model.

I do not have a cause as verifiable as the local one, but I do have a hypothesis with real support in the code: Strands’ Anthropic provider explicitly inserts prompt cache breakpoints (_manage_cache_points in its source) on every request. The LangChain implementation I used for LangGraph does not do that on its own. In a 20-turn conversation that repeats the same system prompt and much of the history on every call, a well-placed cache point saves reprocessing exactly what costs the most. I did not verify this turn by turn against Anthropic’s own cache counters separately; I am stating it as a hypothesis, not a measured fact.

A third data point: the harness you already use every day

Before wrapping up, I added a third point I had not planned: Claude Code, the same CLI I write these posts with, pointed at MiniMax-M3 via cc-minimax (one of my own command-line tools, which only swaps the model’s environment variables). This is not a script of mine: it is real production software, so it did not need a round of audit like the earlier ones, only a spot check against the session’s raw transcript.

Same 20-turn conversation, same gatekeeper, the same SYSTEM_PROMPT appended via --append-system-prompt (without replacing Claude Code’s own, which stays active), same model. The real difference: instead of my 3 toy tools, Claude Code explored the site’s repository with its own native tools (it ended up using Bash in addition to Read/Grep, because --allowedTools with --permission-mode bypassPermissions pre-authorizes without asking but does not restrict which tools are available).

918,648 total tokens, a real $0.87 per the CLI itself. Almost 9 times what LangGraph spent on the same model, and 79 times what Strands Harness spent with temperature matched.

All three harnesses against the same cloud model (MiniMax-M3)total tokens · lower is better
  1. Strands Harness11,618
  2. LangGraph103,952
  3. Claude Code (cc-minimax)918,648

Same 20-turn conversation, same gatekeeper, same system prompt. Claude Code verified against the session's raw transcript, no double-counting.

None of this says Claude Code is inefficient at what it does: it is a general-purpose agent with access to a whole filesystem and a much larger toolset, built for coding tasks, not for answering three questions about a blog. That is exactly the point: using a production hammer to drive a pin costs like a hammer, not like a pin. The right harness depends entirely on the job, and no marketing announcement from any of the three will tell you that this plainly.

What I take away

Whether a framework “saves tokens” or not depends brutally on what you measure it against. Against a local model with a small window and no server-side cache of its own, Strands pays for its own scaffolding (three extra tools, a longer system contract, a summarizer that does not fire in time) and comes out 1.82 times more expensive than the bare option. Against a cloud model with real caching, that same overhead gets diluted and what is left is a real saving, larger than what the announcement itself claims.

And the other thing I take away is about me, not about Strands: the first conclusion I was about to publish was false, and the second was too, halfway. Both times the bug was in my own code, not in the system I was measuring. If this happened with an experiment I designed carefully and audited twice, the lesson is not “trust Strands less”: it is that measuring another framework’s performance, with its hidden defaults (tools you never asked for, contracts you never saw, a counting function you have to use as-is instead of reinventing), is harder than it looks, and publishing without auditing is exactly how numbers that are not real slip through.

How to reproduce it

The four scripts (run_langgraph.py, run_strands.py, run_langgraph_minimax.py, run_strands_minimax.py), the shared module with the real gatekeeper (scope.py from the production chat), and the four raw logs from the final run are documented in this site’s repository. The local server is llama-server (llama.cpp) serving a Qwen3.8-27B Q3_K_M with -c 16384; the cloud model is MiniMax-M3 through its Anthropic-compatible endpoint.

Sources

Comments

No comments yet. The first one is yours.

Reviewed before publishing. The email is not stored and never appears anywhere.