AWS Opens Strands Harness, Says It Beats Claude Code on Cost
Software / news
AWS Opens Strands Harness, Says It Beats Claude Code on Cost
The general-purpose agent runs outside AWS by default and, in AWS's own Terminal-Bench 2.1 tests, cost $56.29 against Claude Code's $248.05 across 89 trials.

Amazon Web Services released Strands Harness on Sept. 21, an open-source, general-purpose AI agent that bundles the tools, context handling and memory a developer would otherwise build by hand, according to the Strands Agents blog. It installs with a single command, pip install strands-harness or npm install @strands-agents/harness, and is licensed under Apache 2.0.
Strands Harness sits on top of the existing Strands Agents SDK, which AWS introduced as an open-source Python framework in May 2025 and later extended to TypeScript. Marc Brooker, a vice president and distinguished engineer at AWS, told The New Stack in a Sept. 21 interview that the SDK alone left developers to solve the same problems repeatedly. "An SDK like the Strands Harness SDK gives you the building blocks, but you still need to decide how to manage context, persist conversations, integrate tools, and guide the agent's behavior," he said. Out of the box, Strands Harness ships with file, shell and web tools, plus built-in handling for context, memory, persistent sessions, prompt caching and delegation to other agents.
Model access defaults to Amazon Bedrock, the only piece of the stack tied to AWS infrastructure by default, Brooker said, but it is "easily overrided to use a different model provider with one line." Strands Harness also supports Anthropic, OpenAI, Google, and locally run models through Ollama or LiteLLM. "Different models have different strengths on reasoning, tool use, and cost," Brooker said. "What doesn't change: context management, sessions, tools, delegation all work the same regardless of provider. No features require Bedrock."
AWS's own benchmark claim, run across six evaluations including GAIA, WebShop and Terminal-Bench 2.1 using the Harbor testing framework on Amazon EC2, is that Strands Harness costs 28% less than comparable harnesses running the same Claude or GPT models, at similar accuracy. That figure is pulled down by one competitor: DeepSeek Harness ran roughly 14% cheaper than Strands Harness on matched runs, and folding it into the wider comparison is what brings the overall savings figure to 28%, AWS said, since DeepSeek Harness also scored lower on every benchmark tested.
The clearest single comparison AWS published is against Claude Code, running Anthropic's Fable 5 model on Terminal-Bench 2.1 across 89 trials.
| Terminal-Bench 2.1, Fable 5 model | Cost | Score |
|---|---|---|
| Strands Harness | $56.29 | 69.7 |
| Claude Code | $248.05 | 61.8 |
| DeepSeek Harness | $40.30 | 59.5 |
On that specific test, Strands Harness cost 77% less than Claude Code while scoring higher, AWS said; DeepSeek Harness undercut Strands Harness on price but trailed both on accuracy. The New Stack's own headline framed the wider comparison as Strands Harness running 45% cheaper than Claude Code and Codex combined, a figure AWS has not broken out on its own blog the way it detailed the Terminal-Bench 2.1 numbers.
AWS also operates Amazon Bedrock AgentCore, a managed service for hosting and operating agents that Brooker said was built by the same team as Strands Harness, though the two live in separate codebases. "AgentCore is an optional hosting layer for teams that want AWS to manage the infrastructure side," he said, adding that Strands Harness can run entirely outside AWS. AWS did not say how many developers have installed Strands Harness since launch, and neither AWS nor The New Stack named a company running it against production traffic as of Sept. 25.
The release adds another vendor-built runtime to a field competing on the same claim: that the reliability work around an agent, not the agent's underlying model, is what teams actually need to buy. It follows the same speed-and-cost framing this site found Inception Labs using for its own model, Mercury 2.5, and AWS's own figures, like the cli-native claims this site checked earlier this month, come from AWS's own test harness rather than an outside benchmark.
Sources
More in Software
- 01postmarketOS Renames Itself Nura After 18-Month SearchThe Linux phone project picked a name tied to 3,000-year-old Sardinian stone towers because regulators would not let it trademark a purely descriptive one.
- 02Paperclip's 90,000 GitHub Stars Rest on One DeveloperAn independent June analysis flagged a nearly 50-to-1 issue-to-contributor ratio behind the fastest-growing agent-orchestration project on GitHub; three months later the backlog has shrunk but the maintainer is still anonymous.
- 03Hindsight Tops GitHub Trending Again, 10 Months After LaunchVectorize's open-source memory layer for AI agents shipped a tokenizer swap and a deadlock fix this month, the kind of unglamorous release cycle that keeps pulling it back to the top of the charts.
- 04Vercel's ScriptC Compiles TypeScript Straight to Native CodeThe experimental compiler skips the JavaScript engine entirely and cut a simple server's cold start from Node's 61.78 milliseconds to 1.78 in an outside benchmark, though a WebKit contributor says its fallback path undercuts the whole idea.