Coding agents do well on small tasks, and lose quality as the tasks grow
Problem 1 of 2, in Tim's words. Below: the outside research behind it, what that research does not show, the questions we still have, and how we plan to measure our answer.
read on 28 Sep 2026 · every outside figure carries its source, its date and whether it was peer-reviewed · BookStack's numbers are featkpr's records of 28 Sep 2026
featkpr holds itnot measured yetnot built yetan idea, not decided
Problem 1 of 2 · Problem 2: teams and testing · Why featkpr, the short version
The evidence #
Research on coding and language-model agents in general. Each figure carries its source, when it was published, and whether it was peer-reviewed.
Success by how long the task takes a person
Issues resolved, by how hard the set is
Five more findings with the same shape #
- 11 of 13 models fell below half their short-context score at 32K tokens of context, on questions that need more than matching words. NoLiMa, ICML 2025 · Feb 2025 · peer-reviewed
- 39% lower the average score when a task arrives over several turns instead of all at once, across six kinds of task and more than 200,000 simulated conversations. Laban et al., 2025 · May 2025 · preprint
- half of 17 models that claim 32K tokens of context or more kept a satisfactory score at 32K. The other half did not. RULER, COLM 2024 · Apr 2024 · peer-reviewed
- each step Accuracy per step falls as the steps add up, even when the model is given the plan and the knowledge. Its own earlier mistakes in the context make the next one more likely. ICLR 2026 · Sep 2025 · peer-reviewed
- the middle What sits in the middle of a long context is used worst. The start and the end are used best, even by models built for long contexts. Lost in the Middle, TACL · Jul 2023 · peer-reviewed
What the evidence does not show #
- None of it measures tests. Every source measures agents on general work: software issues, long texts, conversations. That test quality drops as a flow grows is Tim's experience. We found no study that measures it.
- The line moves. METR finds that the length of task agents can finish doubles about every seven months. A gap measured this year may be smaller next year. METR, NeurIPS 2025 · peer-reviewed
- Some figures come from an interested party. SWE-Bench Pro was written and scored by Scale AI, the company behind the benchmark. Scale AI, 2025 · preprint, vendor
- Snapshot scores age fast. Benchmarks of web agents from 2023 and 2024 showed them far behind people. We leave those scores out, because scores like these change within months.
- A slice does not make a long flow short. Accuracy per step still falls with the number of steps when the plan is given. So the map can keep a task the same size however large the app is; it cannot make one long flow easy. A long flow has to be cut into goals that follow each other. ICLR 2026 · peer-reviewed
- Our own numbers are not an agent's. On BookStack, 497 tests were written from the map by templates, with no model calls. They show the map is enough to write tests from. They say nothing yet about how an agent does with it.
The questions we still have #
- Does an agent given one slice of the map write better tests than the same agent given the whole app? By how much, and at what size of flow does the difference show?
- Does the slice stay small on apps far larger than BookStack, or on programs that are not web apps?
- How long can one flow get before it has to be cut into goals, and does cutting it lose anything?
- As agents improve, does the gap close on its own? If it does, what is the map still for?
- How much of the quality comes from the slice, and how much from the templates that write tests with no model?
- What does a test cost with the slice, and without it?
How the map answers it: our idea, to be measured #
featkpr reads the app once into a map: modules, features, the flows people take and the goals they reach. A task then gets its slice of the map, never the whole app. Our idea is that the slice stays about the same size however large the app grows. Here it is drawn on one real BookStack route.
-
Without the map the whole app
To test one route, an agent reads the whole app and has to keep all of it in mind. BookStack has 312 features in 13 modules. Each dot is one of them.
Entities 83Settings 73Exports 34Activity 28Users 27Uploads 22Access 16Permissions 11Api 8References 4Sorting 4Theming 1Search 1the lit dot: POST /books/{bookSlug}/convert-to-shelf, the route below
-
With the map one route and what it needs
The map hands the agent a slice: this route, the permissions its code checks, and the users who hold them. The other 311 features stay out.
- the route
- POST /books/{bookSlug}/convert-to-shelfTurn a book into a shelf
- its checks
- book-create-all
- book-delete
- book-update
- book-view
- bookshelf-create-all
- its users
- Editor
- Content-only reader
- Reader
- Admin
- a role the test makes
-
One route, five permission checks, 32 tests
featkpr reads the checks from the route's code. A user holds each one or doesn't, so there are 2 × 2 × 2 × 2 × 2 = 32 ways to hold them. featkpr writes one test for each way.
32 tests, one for each way to hold the five checks test 1 2 3 4 5 expects run
aswrong
user1 ✓ E C 2 A 3 A 4 A 5 A 6 A 7 A 8 A 24 more tests, the same pattern
test 1 2 3 4 5 expects run
aswrong
user9 A 10 A 11 A 12 A 13 A 14 A 15 A 16 A 17 A 18 A 19 A 20 A 21 A 22 A 23 A 24 A 25 A 26 A 27 A 28 A 29 A 30 R A 31 A 32 C A holds the permission doesn't hold it ✓expects to get in expects to be kept out E R C run as a user our seed made: the Editor (a BookStack role), the Reader and the Content-only reader (roles featkpr added) run as a role the test makes, holding exactly those (29 of 32) A C the wrong user: the Admin, or the Content-only readerEach column is one test (a row, on a narrow screen). A filled dot means its user holds that permission. Only the outlined test holds all five, so only it expects to get in. The other 31 expect to be kept out.
Every test runs twice. First as its own user, and it must pass. Then as the wrong user, who should get the opposite answer, and now it must fail. A test that still passes wasn't checking the permission. For the outlined test, the wrong user is the Content-only reader, who holds none of the five. For the other 31, it is the Admin, who can do everything.
32 tests, 64 runsThat is for one route. featkpr read 343 routes in BookStack's code. Writing them takes only this route, its five checks and four users. That is the slice. Our idea is that an agent given the slice writes better tests than one given the whole app.
BookStack · featkpr's records of 28 Sep 2026: 312 features, and 32 tests for this route, 32 of them proven · the tests as featkpr's Runs screen shows them on a recorded world, at commit 0f5164ec; no outcomes drawn · routes read in the audit of 26 Sep 2026
Blender · coming: an idea, not built
A 3D creation suite: modelling, animation, rendering and video editing. Open source, and a desktop program, not a web app.
Far larger than BookStack, with no routes or pages to crawl. featkpr would need a reader for its code and a way to walk a desktop program's screens. It would show whether one route's slice stays small on a product of another size and shape.
featkpr has not read Blender, so there is nothing to draw. Tim, raised 28 Sep · the idea on the board
LibreOffice · coming: an idea, not built
An office suite: documents, spreadsheets, slides, drawings and databases. Open source, and a desktop program, not a web app.
Also far larger than BookStack, and several programs built from one code base. It would show whether a slice stays small when many features share the same code.
featkpr has not read LibreOffice, so there is nothing to draw. Tim, raised 28 Sep · the idea on the board
Not measured yet. On BookStack the map exists and 497 tests were written from it, from templates, with no model calls. No run compares an agent with the map against an agent alone. The comparison is designed (40 scenarios in four sizes, the same agent with and without the slice); its chart goes here only when it has numbers.
How we will measure it #
a proposal from our research of 28 Sep 2026, not on the plan: on 28 Sep Tim decided not to run it yet · paid model runs need a capped key and Tim's approval
The same agent writes tests for the same scenarios twice: once on its own, with the code and a running copy, and once with the map's slice. The scenarios come from BookStack's stored goals and flows, in four sizes, ten of each.
Four sizes of scenario, fixed before any run #
| size | screens | steps | inputs | roles | features involved |
|---|---|---|---|---|---|
| C1 | 1 | up to 3 | up to 1 | 1 | 1 |
| C2 | 2 to 3 | 4 to 8 | 2 to 5 | 1 | 1 |
| C3 | 4 to 6 | 9 to 15 | 6 to 12 | 1 to 2 | 2, e.g. a page needs a book and a chapter |
| C4 | 7 or more | 16 or more | 13 or more | 2 or more, with a permission change midway | 3 or more |
- C1
- C2
- C3
- C4
Three ways to write each test #
- An agent alone a coding agent with the repository and a throwaway running copy, given the scenario in plain words
- The same agent with the slice the same prompt, plus the feature, the flow's screens and routes, its permission checks, the right and the wrong user, and the records it needs
- featkpr's templates no model at all, where a template covers the scenario: the reference line
What counts as a good test, in order #
- It runs.
- It passes as the right user.
- It fails as the wrong user.
- It checks the goal's outcome, not just a page loading.
- It goes red when a fault is planted in the app.
- A person, not told which way it was written, scores it 0 to 3.
- It gives the same answer over five reruns.
The rule, written down before any run #
The idea holds if tests written with the slice lose no more than 10 points, from the smallest size to the largest, on proven and fault caught together, while tests written alone lose at least 25. Otherwise this page says what was measured instead. Then the same again on Shlink, the second app, before any claim.
A pilot first: three scenarios of each size, one run each, for both agent arms. That is 24 runs on a key capped at $25, to learn what the agent alone costs before the full run is asked for.
What the plan already holds for it #
- How complex each goal is, drafted from what drives it; a person's word is feedback on the draft shipped 27 Sep
- Examples from much larger products: Blender, LibreOffice an idea, Tim, raised 28 Sep
Sources #
each read on 28 Sep 2026; the date is when it was published, the mark says how it was checked
- Kwa, West, Becker et al. (METR), Measuring AI Ability to Complete Long Software Tasks, NeurIPS 2025. The two figures as METR's post of 19 Mar 2025 states them; the doubling time, about every 7 months, is the paper's. 19 Mar 2025 · peer-reviewed arxiv.org/abs/2503.14499 · metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- Deng, Da, Pan et al. (Scale AI), SWE-Bench Pro, arXiv 2509.16941, Sep 2025. Written by the benchmark's own vendor; the paper's framing of Opus 4.1 and GPT-5. Sep 2025 · preprint, vendor arxiv.org/abs/2509.16941
- Modarressi et al. (Adobe), NoLiMa: Long-Context Evaluation Beyond Literal Matching, arXiv 2502.05167 (Feb 2025), ICML 2025. Feb 2025 · peer-reviewed arxiv.org/abs/2502.05167
- Laban, Hayashi, Zhou, Neville, LLMs Get Lost In Multi-Turn Conversation, arXiv 2505.06120, May 2025. May 2025 · preprint arxiv.org/abs/2505.06120
- Hsieh et al. (NVIDIA), RULER: What's the Real Context Size of Your Long-Context Language Models?, arXiv 2404.06654 (Apr 2024), COLM 2024. Apr 2024 · peer-reviewed arxiv.org/abs/2404.06654
- Sinha, Arun, Goel, Staab, Geiping, The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs, arXiv 2509.09677 (Sep 2025), ICLR 2026. Sep 2025 · peer-reviewed arxiv.org/abs/2509.09677
- Liu, Lin, Hewitt et al., Lost in the Middle: How Language Models Use Long Contexts, arXiv 2307.03172 (Jul 2023), TACL. Jul 2023 · peer-reviewed arxiv.org/abs/2307.03172
See the map on BookStack #
A call of about 30 minutes: the map, a flow waiting on a verdict, and a route's slice. Problem 2: teams and testing.