# Coding agents do well on small tasks, and lose quality as the tasks grow

Problem 1 of 2, in Tim's words. Below: the outside research behind it, what that research does not show, the questions we still have, and how we plan to measure our answer.

read on 28 Sep 2026 · every outside figure carries its source, its date and whether it was peer-reviewed · BookStack's numbers are featkpr's records of 28 Sep 2026

featkpr holds it not measured yet not built yet an idea, not decided

Problem 1 of 2 · [Problem 2: teams and testing](https://featkpr.com/why/teams) · [Why featkpr, the short version](https://featkpr.com/why)

## The evidence

Research on coding and language-model agents in general. Each figure carries its source, when it was published, and whether it was peer-reviewed.

Success by how long the task takes a person

under 4 minutes **almost 100%** more than about 4 hours **under 10%**

[METR, NeurIPS 2025](https://arxiv.org/abs/2503.14499) · peer-reviewed

Issues resolved, by how hard the set is

SWE-Bench Verified **over 70%** SWE-Bench Pro: fixes average 107 lines across 4 files **about 23%**

[Scale AI, 2025](https://arxiv.org/abs/2509.16941) · preprint, vendor

### Five more findings with the same shape

- **11 of 13** models fell below half their short-context score at 32K tokens of context, on questions that need more than matching words. [NoLiMa, ICML 2025](https://arxiv.org/abs/2502.05167) · Feb 2025 · peer-reviewed
- **39% lower** the average score when a task arrives over several turns instead of all at once, across six kinds of task and more than 200,000 simulated conversations. [Laban et al., 2025](https://arxiv.org/abs/2505.06120) · May 2025 · preprint
- **half of 17** models that claim 32K tokens of context or more kept a satisfactory score at 32K. The other half did not. [RULER, COLM 2024](https://arxiv.org/abs/2404.06654) · Apr 2024 · peer-reviewed
- **each step** Accuracy per step falls as the steps add up, even when the model is given the plan and the knowledge. Its own earlier mistakes in the context make the next one more likely. [ICLR 2026](https://arxiv.org/abs/2509.09677) · Sep 2025 · peer-reviewed
- **the middle** What sits in the middle of a long context is used worst. The start and the end are used best, even by models built for long contexts. [Lost in the Middle, TACL](https://arxiv.org/abs/2307.03172) · Jul 2023 · peer-reviewed

## What the evidence does not show

- **None of it measures tests.** Every source measures agents on general work: software issues, long texts, conversations. That test quality drops as a flow grows is Tim's experience. We found no study that measures it.
- **The line moves.** METR finds that the length of task agents can finish doubles about every seven months. A gap measured this year may be smaller next year. [METR, NeurIPS 2025](https://arxiv.org/abs/2503.14499) · peer-reviewed
- **Some figures come from an interested party.** SWE-Bench Pro was written and scored by Scale AI, the company behind the benchmark. [Scale AI, 2025](https://arxiv.org/abs/2509.16941) · preprint, vendor
- **Snapshot scores age fast.** Benchmarks of web agents from 2023 and 2024 showed them far behind people. We leave those scores out, because scores like these change within months.
- **A slice does not make a long flow short.** Accuracy per step still falls with the number of steps when the plan is given. So the map can keep a task the same size however large the app is; it cannot make one long flow easy. A long flow has to be cut into goals that follow each other. [ICLR 2026](https://arxiv.org/abs/2509.09677) · peer-reviewed
- **Our own numbers are not an agent's.** On BookStack, 497 tests were written from the map by templates, with no model calls. They show the map is enough to write tests from. They say nothing yet about how an agent does with it.

## The questions we still have

- Does an agent given one slice of the map write better tests than the same agent given the whole app? By how much, and at what size of flow does the difference show?
- Does the slice stay small on apps far larger than BookStack, or on programs that are not web apps?
- How long can one flow get before it has to be cut into goals, and does cutting it lose anything?
- As agents improve, does the gap close on its own? If it does, what is the map still for?
- How much of the quality comes from the slice, and how much from the templates that write tests with no model?
- What does a test cost with the slice, and without it?

## How the map answers it: our idea, to be measured

featkpr reads the app once into a map: modules, features, the flows people take and the goals they reach. A task then gets its slice of the map, never the whole app. Our idea is that the slice stays about the same size however large the app grows. Here it is drawn on one real BookStack route.

Entities 83 Settings 73 Exports 34 Activity 28 Users 27 Uploads 22 Access 16 Permissions 11 Api 8 References 4 Sorting 4 Theming 1 Search 1

the lit dot: POST /books/{bookSlug}/convert-to-shelf, the route below

**With the map** one route and what it needs

The map hands the agent a slice: this route, the permissions its code checks, and the users who hold them. The other 311 features stay out.

**the route**

**POST** /books/{bookSlug}/convert-to-shelf Turn a book into a shelf

**its checks**

- book-create-all
- book-delete
- book-update
- book-view
- bookshelf-create-all

**its users**

- Editor
- Content-only reader
- Reader
- Admin
- a role the test makes

#### One route, five permission checks, 32 tests

featkpr reads the checks from the route's code. A user holds each one or doesn't, so there are 2 × 2 × 2 × 2 × 2 = 32 ways to hold them. featkpr writes one test for each way.

book-delete book-update book-view bookshelf-create-all expects run as wrong user

- **✓** E C
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- A
- R A
- A
- C A
- book-create-all
- **2** book-delete
- **3** book-update
- **4** book-view
- **5** bookshelf-create-all

32 tests, one for each way to hold the five checks

| test | 1 | 2 | 3 | 4 | 5 | expects | run as | wrong user |

|---|---|---|---|---|---|---|---|---|

| 1 |  |  |  |  |  | **✓** | **E** | **C** |

| 2 |  |  |  |  |  |  |  | **A** |

| 3 |  |  |  |  |  |  |  | **A** |

| 4 |  |  |  |  |  |  |  | **A** |

| 5 |  |  |  |  |  |  |  | **A** |

| 6 |  |  |  |  |  |  |  | **A** |

| 7 |  |  |  |  |  |  |  | **A** |

| 8 |  |  |  |  |  |  |  | **A** |

24 more tests, the same pattern

| test | 1 | 2 | 3 | 4 | 5 | expects | run as | wrong user |

|---|---|---|---|---|---|---|---|---|

| 9 |  |  |  |  |  |  |  | **A** |

| 10 |  |  |  |  |  |  |  | **A** |

| 11 |  |  |  |  |  |  |  | **A** |

| 12 |  |  |  |  |  |  |  | **A** |

| 13 |  |  |  |  |  |  |  | **A** |

| 14 |  |  |  |  |  |  |  | **A** |

| 15 |  |  |  |  |  |  |  | **A** |

| 16 |  |  |  |  |  |  |  | **A** |

| 17 |  |  |  |  |  |  |  | **A** |

| 18 |  |  |  |  |  |  |  | **A** |

| 19 |  |  |  |  |  |  |  | **A** |

| 20 |  |  |  |  |  |  |  | **A** |

| 21 |  |  |  |  |  |  |  | **A** |

| 22 |  |  |  |  |  |  |  | **A** |

| 23 |  |  |  |  |  |  |  | **A** |

| 24 |  |  |  |  |  |  |  | **A** |

| 25 |  |  |  |  |  |  |  | **A** |

| 26 |  |  |  |  |  |  |  | **A** |

| 27 |  |  |  |  |  |  |  | **A** |

| 28 |  |  |  |  |  |  |  | **A** |

| 29 |  |  |  |  |  |  |  | **A** |

| 30 |  |  |  |  |  |  | **R** | **A** |

| 31 |  |  |  |  |  |  |  | **A** |

| 32 |  |  |  |  |  |  | **C** | **A** |

holds the permission doesn't hold it **✓** expects to get in expects to be kept out **E** **R** **C** run as a user our seed made: the Editor (a BookStack role), the Reader and the Content-only reader (roles featkpr added) run as a role the test makes, holding exactly those (29 of 32) **A** **C** the wrong user: the Admin, or the Content-only reader

Each column is one test (a row, on a narrow screen). A filled dot means its user holds that permission. **Only the outlined test holds all five, so only it expects to get in. The other 31 expect to be kept out.**

Every test runs twice. First as its own user, and it must pass. Then as the wrong user, who should get the opposite answer, and now it must fail. A test that still passes wasn't checking the permission. For the outlined test, the wrong user is the Content-only reader, who holds none of the five. For the other 31, it is the Admin, who can do everything.

**32 tests, 64 runs**

That is for one route. featkpr read 343 routes in BookStack's code. Writing them takes only this route, its five checks and four users. That is the slice. Our idea is that an agent given the slice writes better tests than one given the whole app.

BookStack · featkpr's records of 28 Sep 2026: 312 features, and 32 tests for this route, 32 of them proven · the tests as featkpr's Runs screen shows them on a recorded world, at commit 0f5164ec; no outcomes drawn · routes read in the audit of 26 Sep 2026

**Blender** · coming: an idea, not built

A 3D creation suite: modelling, animation, rendering and video editing. Open source, and a desktop program, not a web app.

Far larger than BookStack, with no routes or pages to crawl. featkpr would need a reader for its code and a way to walk a desktop program's screens. It would show whether one route's slice stays small on a product of another size and shape.

featkpr has not read Blender, so there is nothing to draw. Tim, raised 28 Sep · [the idea on the board](https://featkpr.com/roadmap#larger-examples)

**LibreOffice** · coming: an idea, not built

An office suite: documents, spreadsheets, slides, drawings and databases. Open source, and a desktop program, not a web app.

Also far larger than BookStack, and several programs built from one code base. It would show whether a slice stays small when many features share the same code.

featkpr has not read LibreOffice, so there is nothing to draw. Tim, raised 28 Sep · [the idea on the board](https://featkpr.com/roadmap#larger-examples)

Not measured yet. On BookStack the map exists and 497 tests were written from it, from templates, with no model calls. No run compares an agent with the map against an agent alone. The comparison is designed (40 scenarios in four sizes, the same agent with and without the slice); its chart goes here only when it has numbers.

## How we will measure it

a proposal from our research of 28 Sep 2026, not on the plan: on 28 Sep Tim decided not to run it yet · paid model runs need a capped key and Tim's approval

The same agent writes tests for the same scenarios twice: once on its own, with the code and a running copy, and once with the map's slice. The scenarios come from BookStack's stored goals and flows, in four sizes, ten of each.

### Four sizes of scenario, fixed before any run

| size | screens | steps | inputs | roles | features involved |

|---|---|---|---|---|---|

| C1 | 1 | up to 3 | up to 1 | 1 | 1 |

| C2 | 2 to 3 | 4 to 8 | 2 to 5 | 1 | 1 |

| C3 | 4 to 6 | 9 to 15 | 6 to 12 | 1 to 2 | 2, e.g. a page needs a book and a chapter |

| C4 | 7 or more | 16 or more | 13 or more | 2 or more, with a permission change midway | 3 or more |

tests proven, and fault caught not measured: nothing is drawn until it runs

- C1
- C2
- C3
- C4

When it runs, the two measured lines go here, with the number of scenarios, the date, the model and the intervals. Until then, nothing is drawn.

### Three ways to write each test

- **An agent alone** a coding agent with the repository and a throwaway running copy, given the scenario in plain words
- **The same agent with the slice** the same prompt, plus the feature, the flow's screens and routes, its permission checks, the right and the wrong user, and the records it needs
- **featkpr's templates** no model at all, where a template covers the scenario: the reference line

### What counts as a good test, in order

- It runs.
- It passes as the right user.
- It fails as the wrong user.
- It checks the goal's outcome, not just a page loading.
- It goes red when a fault is planted in the app.
- A person, not told which way it was written, scores it 0 to 3.
- It gives the same answer over five reruns.

### The rule, written down before any run

The idea holds if tests written with the slice lose no more than 10 points, from the smallest size to the largest, on proven and fault caught together, while tests written alone lose at least 25. Otherwise this page says what was measured instead. Then the same again on Shlink, the second app, before any claim.

A pilot first: three scenarios of each size, one run each, for both agent arms. That is 24 runs on a key capped at $25, to learn what the agent alone costs before the full run is asked for.

### What the plan already holds for it

- [How complex each goal is, drafted from what drives it; a person's word is feedback on the draft](https://featkpr.com/roadmap#goals-with-proof) shipped 27 Sep
- [Examples from much larger products: Blender, LibreOffice](https://featkpr.com/roadmap#larger-examples) an idea, Tim, raised 28 Sep

## Sources

each read on 28 Sep 2026; the date is when it was published, the mark says how it was checked

- Kwa, West, Becker et al. (METR), Measuring AI Ability to Complete Long Software Tasks, NeurIPS 2025. The two figures as METR's post of 19 Mar 2025 states them; the doubling time, about every 7 months, is the paper's. 19 Mar 2025 · peer-reviewed [arxiv.org/abs/2503.14499](https://arxiv.org/abs/2503.14499) · [metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)
- Deng, Da, Pan et al. (Scale AI), SWE-Bench Pro, arXiv 2509.16941, Sep 2025. Written by the benchmark's own vendor; the paper's framing of Opus 4.1 and GPT-5. Sep 2025 · preprint, vendor [arxiv.org/abs/2509.16941](https://arxiv.org/abs/2509.16941)
- Modarressi et al. (Adobe), NoLiMa: Long-Context Evaluation Beyond Literal Matching, arXiv 2502.05167 (Feb 2025), ICML 2025. Feb 2025 · peer-reviewed [arxiv.org/abs/2502.05167](https://arxiv.org/abs/2502.05167)
- Laban, Hayashi, Zhou, Neville, LLMs Get Lost In Multi-Turn Conversation, arXiv 2505.06120, May 2025. May 2025 · preprint [arxiv.org/abs/2505.06120](https://arxiv.org/abs/2505.06120)
- Hsieh et al. (NVIDIA), RULER: What's the Real Context Size of Your Long-Context Language Models?, arXiv 2404.06654 (Apr 2024), COLM 2024. Apr 2024 · peer-reviewed [arxiv.org/abs/2404.06654](https://arxiv.org/abs/2404.06654)
- Sinha, Arun, Goel, Staab, Geiping, The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs, arXiv 2509.09677 (Sep 2025), ICLR 2026. Sep 2025 · peer-reviewed [arxiv.org/abs/2509.09677](https://arxiv.org/abs/2509.09677)
- Liu, Lin, Hewitt et al., Lost in the Middle: How Language Models Use Long Contexts, arXiv 2307.03172 (Jul 2023), TACL. Jul 2023 · peer-reviewed [arxiv.org/abs/2307.03172](https://arxiv.org/abs/2307.03172)

## See the map on BookStack

A call of about 30 minutes: the map, a flow waiting on a verdict, and a route's slice. [Problem 2: teams and testing](https://featkpr.com/why/teams) .

[Ask for a private demo](https://featkpr.com/demo)

---

The page this twin stands for: https://featkpr.com/why/agents. Every page on this site has a `.md` twin, and answers `Accept: text/markdown`.
