Ponytail Has 154,000 Stars for a Seven-Step Ruleset Whose Benchmark Is Four Runs on One Repo
Software / news
Ponytail Has 154,000 Stars for a Seven-Step Ruleset Whose Benchmark Is Four Runs on One Repo
The MIT-licensed rule file tells coding agents to write less code. Its 54% cut in lines comes from the author's own test with Claude Haiku 4.5.
Ponytail, a rule file that tells coding agents to write as little code as possible, gained 1,894 stars in a day on GitHub trending on October 4, 2026, and sits at 154,408 stars. Its published evidence is a benchmark of twelve feature tasks on one FastAPI and React repository, run with Claude Haiku 4.5, with n=4.
The repository, DietrichGebert/ponytail, is owned by an individual account, carries an MIT licence and was created on June 12, 2026, according to GitHub's API. The same API lists 227 open issues. The README page lists 55, so the count depends on where you read it.
The seven rungs a model is told to climb
Ponytail is not code. It is a ruleset that an agent reads before writing, a decision ladder applied after it understands the problem:
- Does this need to exist? If not, skip it.
- Is it already in the codebase? Reuse it.
- Does the standard library do it? Use that.
- Is there a native platform feature? Use it.
- Does an installed dependency do it? Use it.
- Does it fit in one line? Write one line.
- Otherwise, write the minimum that works.
The README says "validation, data-loss handling, security, and accessibility are never on the chopping block." It lists plugins or instruction files for more than 20 agents, among them Claude Code, Cursor, GitHub Copilot, Codex and Gemini CLI.
What the benchmark measured
The README reports reductions against a no-rule baseline, from "real agent sessions" on the twelve tasks. Every figure is the author's own.
| Metric | Reduction with ponytail |
|---|---|
| Lines of code | 54% (up to 94%) |
| Time | 27% |
| Tokens | 22% |
| Cost | 20% |
- Lines of code54 %
- Time27 %
- Tokens22 %
- Cost20 %
Source: ponytail README, author-supplied, accessed 2026-10-04
The README's best example is a date picker that fell from 404 lines to 23. It says the gains are largest where an agent over-engineers and small where the code is already tight.
It also claims "100% safe" and says ponytail is "the only arm that cuts every metric, and the only one that stays fully safe while doing it." The word arm implies other conditions were compared, and the page does not say what they were or how safety was scored.
What four runs cannot show
Four runs per task on a single repository is a small sample. Haiku 4.5 is a small model, and an instruction to write less may behave differently on a larger one. Counting lines rewards terse code, not code that is easier to maintain. A 23-line date picker is shorter than a 404-line one, and nothing in the README shows the two do the same job.
Stars are not adoption either. A repository gaining 1,894 in a day tells you how many people clicked, not how many kept the rule file installed after the first week.
The cost saving is the figure to check against your own bill. Whether a 20% drop survives on a repository with a different shape is untested.
Where it sits among agent skill packs
Ponytail belongs to a growing shelf of rule and skill packs that spread by star count. We have also covered Agent-Reach and Dwarfstar, two other repositories that collected stars quickly.
The practical test is cheap. Install the rule on a branch, give the agent two tasks you have already merged, and compare diffs and token counts. The author's repository is the only source of the 54% figure until someone else publishes a reproduction.
Sources
More in Software
- 01OpenCut Has 92,200 Stars, but the Editor People Use Is the Classic One and the Rewrite Is Not Taking ContributionsThe open-source CapCut alternative rebuilt its default branch in May. The README and a third-party walkthrough disagree on how much of the new code is Rust.
- 02Impeccable's Design Detector Runs Without a Model, but Its Open Issues Show Gaps Outside .htmlPaul Bakaus's design skill for coding agents ships 61 deterministic rules you can run from the command line. The bug tracker says where they are least reliable.
- 03A Hacker News Post Says Agents Need Documentation, Not Memory, and Its Author Wrote the Plugin That Does ThatKevin Liao's October 3 essay attacks snippet-recall memory plugins and promotes Operator Memory. A separate September essay argues the real gap is neither memory nor documents.
- 04Agent Reach, at 90,900 Stars, Reads X and Reddit for Your Agent Through Your Own CookiesThe MIT-licensed CLI routes agents to 20-plus sites with a backup backend per channel. Its README admits the login channels can get an account banned.