TesterArmy's e2e Gained 1,398 Stars in a Day. Its Docs Publish No Accuracy Figures
Software / news
TesterArmy's e2e Gained 1,398 Stars in a Day. Its Docs Publish No Accuracy Figures
The Apache-2.0 framework lets an agent drive an app from a plain-English goal, then replays the recorded steps. The one independent benchmark of its decision-model option tested a different job.

TesterArmy's open-source end-to-end testing framework, e2e, gained 1,398 GitHub stars in one day to reach about 4,900, according to the day's trending list and the repository's API. It is a test runner where a step like agent.act('upgrade the workspace to the Pro plan') hands a goal to a model, followed by ordinary assertions.
The repository was created on July 22, 2026, is licensed Apache-2.0, and its npm package e2e is at version 0.17.0, published October 4. The README says it is "in active development on the way to 1.0" and that "APIs and config can still change between minor releases."

What a test looks like and what it costs to rerun
The README's example opens a billing page, tells the agent to upgrade to the Pro plan, asks it to assert that the invoice preview shows a prorated amount, then checks a status element with a normal locator. A step that a later assertion verifies "records its actions, and the next run replays them with no model calls until the app changes." Tests without agent steps need no model.
The quickstart lists Node.js 24.8 or newer, WSL on Windows, and Xcode or the Android SDK for mobile. It names more than 50 model providers, from ChatGPT Plus to Ollama, and gives no pricing, because cost is whatever the chosen provider charges. It also does not say what happens when an app changes mid-run beyond a --no-cache flag that skips replay.
The CLI sends anonymous usage data by default, such as which commands run and where runs fail, with no test content. It is switched off with npx e2e telemetry disable or E2E_TELEMETRY_DISABLED=1.
The decision-model option
The package @e2e-dev/decision is the cost-saving path. Its documentation describes running "bounded actions with a decision model and a small text model instead of a full LLM agent." The decision model picks an operation and a target element from the page's semantic tree, and a small text model writes a value only when the choice is type. The docs give jev-latest as the decision model and inception/mercury-2.5 as the text model.
The documentation "provides no metrics" on latency, cost or accuracy: the page The Terminal read contains none, and it names no failure rate for a wrong element choice.
What the one independent benchmark measured
On October 2, Red Hat Developer published a benchmark of decision models against classifiers and LLM judges by Dr. Rob Geada, Dr. Mac Misiura and Shelton Cyril. It compared nine approaches on prompt-injection detection and content safety, including Jev, which the article attributes to TypeSafe. That is a different task from choosing a button, so it says nothing direct about e2e's reliability.
- Qwen3.6-35B (LLM judge)89.31 %
- deberta-v3-base (200M classifier)89.01 %
- DiffusionGemma87.72 %
- Jev86.35 %
Source: Red Hat Developer, October 2, 2026, accessed 2026-10-06
The authors concluded that "decision models like Jev do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy." Median latency on prompt injection was 348.1 ms for Jev against 54.1 ms for the 200-million-parameter classifier. On content safety Jev ranked first at 86.20%. The authors caution that hardware differed between models and that remote calls added about 56 ms from the UK to the US.
The objection from the launch thread
TesterArmy's hosted product was introduced on Hacker News on June 18, 2026, drawing 132 points. One commenter asked why anyone needs an outside tool if Opus already writes the code and "would know best what E2E tests to write." The poster, user okwasniewski, replied that static tests are brittle because they "rely on selectors, need wait times" and struggle with dynamic content.
For what unsupervised agents have done to a live site, see Wikimedia's report on OpenAI agents editing its wikis, and for a model release with the same gap between claim and measurement, Runway's Praxis-1. The next marker here is the 1.0 release, which the README says is coming and does not date.
Sources
More in Software
- 01VB6 Studio Web 0.6.0 Rebuilds Visual Basic 6 in a Browser, Minus COM and OCXWieslaw Soltes's MIT-licensed project reads .vbp project files and exports standalone HTML, but its README rules out the native components that classic VB6 programs lean on.
- 02REA Hands Coding Agents a Disassembler, and 2,963 Stars Followed in a DayThe MIT-licensed toolkit from a developer called morluto wires Ghidra, Hopper and a JavaScript analyser into agents over MCP, and its README pitches rebuilding a feature you have only seen running.
- 03Polars 2.0 Makes Streaming the Default and Stops Promising Row OrderThe release, announced Oct. 6 by creator Ritchie Vink, adds no headline feature, but a join or group-by that used to return rows in input order may now return them in any order.
- 04Cloudflare's Web Search API Passes Through Exa, Linkup and Ceramic.ai PricesThe beta, announced Oct. 2 through AI Gateway, adds no markup, so the bill depends on which provider string a developer types and the launch post prints no price.