# Nobody holds a map of what your app does, or how its features depend on each other.

The problem has two sides. Coding agents lose quality as tasks grow, and teams usually don't like testing. featkpr reads the code and a running copy, draws the features, the flows they form and the goals they reach, and puts each one to a person to keep or drop. Tests, and anything else, sit on top of that map.

as of 28 Sep 2026 · every outside figure carries its source and whether it was peer-reviewed · featkpr's app screens run on sample data; BookStack's numbers are its own

featkpr holds it waiting on a person, or not measured yet not built yet an idea, not decided

## Coding agents do well on small tasks, and lose quality as the tasks grow

problem 1 of 2

Ask a coding agent for a test of a short flow and the test is usually good. Make the flow longer, across more screens, roles and features, and the tests get worse. Research on agents in general measures the same shape:

Success by how long the task takes a person

under 4 minutes **almost 100%** more than about 4 hours **under 10%**

[METR, NeurIPS 2025](https://arxiv.org/abs/2503.14499) · peer-reviewed

Issues resolved, by how hard the set is

SWE-Bench Verified **over 70%** SWE-Bench Pro: fixes average 107 lines across 4 files **about 23%**

[Scale AI, 2025](https://arxiv.org/abs/2509.16941) · preprint, vendor

None of this research measures tests directly. Our idea, to be measured: the map hands each task one slice of the app, so the task stays about the same size however large the app grows.

[The research behind it, and what it does not show](https://featkpr.com/why/agents)

## Teams usually don't like testing

problem 2 of 2

Developers, owners and even QA. Suites rarely keep up with new features, what they catch is often basic, and the fix holds the release. That is Tim's view. Outside research supports part of it:

- **47%** of about 3,700 people in GitLab's 2020 survey named testing the top cause of release delays (49% the year before) [GitLab survey, 2020](https://about.gitlab.com/blog/devsecops-survey-released/) · 18 May 2020 · vendor survey
- **58%** of 284 developers prefer writing code to writing tests [Straubinger and Fraser, ISSRE 2023](https://arxiv.org/abs/2309.01154) · Sep 2023 · peer-reviewed
- **84%** of the changes from pass to fail seen at Google involved a flaky test [Micco, Google, 2016](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html) · 27 May 2016 · engineering blog, Google's own data

Each bar is a share of its own group, so the bars don't compare with each other.

No survey we found asks owners, nothing shows that what tests catch is basic, and a 2025 QA survey points the other way. Our idea, to be measured: tests drafted from a map the team approved keep pace with new features.

[The research behind it, and what it does not show](https://featkpr.com/why/teams)

## The idea: one map between the code and the running app

Docs describe what the app should do, and they drift. Tests check the running app, one screen at a time. featkpr keeps the layers between: the features, how they lead into each other, the flows they form and the goals they reach. Each piece goes to a person to keep or drop.

93% of the people asked had run into incomplete or outdated docs. [GitHub, 2017](https://opensourcesurvey.org/2017/) · survey

**The running app** what people see and click

docs describe it · tests check it · featkpr crawls a throwaway copy

**Flows to goals** the paths people take to reach a goal, screen by screen

**How features relate** which feature leads to which, and the flows that use each link; the data they share is being built

**featkpr's map**

**Features, by module** what the app can do, each named

**The code** routes, handlers, permission checks, templates

featkpr reads it

Docs and tests attach at the running end. featkpr reads the code and a running copy, and keeps the layers between as the map. Screens of featkpr's app run on sample data.

### How the map answers each side: our idea, to be measured

- **Coding agents.** A task gets one slice of the map, never the whole app, so it stays about the same size as the app grows. [The idea in full](https://featkpr.com/why/agents#answer)
- **Teams and testing.** The team approves the map one question at a time. Tests are drafted from the approved map, so a new feature reaches the map first and its tests follow. [The idea in full](https://featkpr.com/why/teams#answer)

## What featkpr drew on BookStack, and what it missed

BookStack is an open-source wiki we don't own. featkpr read its code and a running copy, with nobody from BookStack helping. Then it started again with no help at all and counted what it missed.

- **312** features in 13 modules, named from the code featkpr's records · 28 Sep 2026
- **167** screens crawled on a running copy, signed in as Admin the crawl · 25 Sep 2026 · commit 0f5164ec
- **149** flows drafted to 88 goals, each waiting on a person the Flows board · 28 Sep 2026
- **not held yet** how features lead into each other: the Flows tree, each link with the flows that use it the real app's Flows tree · 28 Sep 2026
- **not held yet** the data features share being built

605 of 2,372 items missed, by reason

- no detector reads it yet **295**
- cut by a limit on the run **199**
- the step that finds it isn't built **67**
- needed a record that wasn't seeded **28**
- needed a role the crawl didn't use **10**
- no page links to it **3**
- stopped by a safety guard **3**

BookStack · the fresh-start audit of 26 Sep 2026 · commit 0f5164ec · 592 of them featkpr's own gaps · [BookStack in the library](https://featkpr.com/library/bookstack)

## Map first, then anything on top

The idea since 8 Sep: read the app once into a map, have people check it, and build everything else from the checked map. What sits on top today is the start; the slots above it come from the plan.

- [**Tests per route and role, each sent again as the wrong user** works today](https://featkpr.com/roadmap#tests-wrong-user-check)
- [**What a change reaches, and a verdict on it** being built](https://featkpr.com/roadmap#merge-request-verdict)
- [**What a test changed, and OpenTelemetry on runs** being built](https://featkpr.com/roadmap#effects-on-runs)
- [**Every state change traced, and your own telemetry** from 19 Oct](https://featkpr.com/roadmap#whole-system-tracing)
- [**Flows run as browser tests** decided, no week](https://featkpr.com/roadmap#browser-tests)
- [**Tests in TypeScript** decided, no week](https://featkpr.com/roadmap#typescript-tests)
- [**Selenium and Cucumber outputs** decided, no week](https://featkpr.com/roadmap#selenium-cucumber)
- [**The flags and errors you already track** decided, no week](https://featkpr.com/roadmap#flags-and-errors)
- [**Performance budgets per goal** an idea](https://featkpr.com/roadmap#performance-budgets)
- [**Load tests** decided, no week](https://featkpr.com/roadmap#load-tests-k6)
- [**The map, served to coding agents** an idea](https://featkpr.com/roadmap#agents-server)
- [**A guide to your product** an idea](https://featkpr.com/roadmap#product-guide)

**Your review: keep, drop or rewrite each piece, one question at a time**

**The map: modules, features, how they lead into each other, the flows they form and the goals they reach**

**The data features share** being built

**What featkpr reads: your code and a test copy that says how closely it copies real hosting**

On top, from the plan of 28 Sep 2026: a filled node works today, a ring is being built, a hollow node is planned, a dotted one is an idea. Each links to its card.

## What other tools map, and who decides

AI test tools map an app so they can write tests, and mostly decide what goes in themselves. The nearest, Momentic's app graph, lets you approve each addition; its docs say it is in alpha. Code maps read the code, and map functions and files. Analytics draw flows from real traffic, so a path nobody took isn't there. Security testers check who may reach what, API by API, without a map of features. [9](https://featkpr.com/compare#s-9) [35](https://featkpr.com/compare#s-35) [44](https://featkpr.com/compare#s-44) [56](https://featkpr.com/compare#s-56)

keeps: features or flows decides: a person approves each piece

AI test tools: Momentic (alpha)

keeps: features, flows and goals, linked decides: a person approves each piece

**featkpr**

keeps: tests or recorded sessions decides: people steer or edit it

AI test tools: Playwright

keeps: code, pages, APIs or events decides: people steer or edit it

Code maps: Swimm

keeps: features or flows decides: people steer or edit it

3 AI test tools

Analytics: Fullstory

Research: ScenGen

keeps: tests or recorded sessions decides: the tool decides

3 AI test tools

keeps: code, pages, APIs or events decides: the tool decides

AI test tools: mabl, Qodo

4 code maps

5 security testers

keeps: features or flows decides: the tool decides

AI test tools: Octomind

Analytics: PostHog, Heap

Research: AutoE2E and E2EBench

Each product sits where its own pages put it; the full comparison quotes them, cell by cell. featkpr's dot is open: on BookStack its flows still wait on a person, 149 of 149 (28 Sep 2026).

[The full comparison](https://featkpr.com/compare)

The closest work we found is AutoE2E (UBC, 2024, a preprint): it infers a web app's features and writes end-to-end tests from them. We found none that keeps a checked map, with relations and a person deciding each piece. [AutoE2E, 2024](https://arxiv.org/abs/2408.01894) · preprint

## Our idea, to be measured

What BookStack shows today, and what nobody has measured yet. Each problem page says how we plan to measure it.

Shown on BookStack BookStack · featkpr's records of 28 Sep 2026

**The map** 312 features in 13 modules, and 149 flows drafted **Tests from the map** 497 written by templates, with no model calls; 311 ran on 28 Sep 2026, 06:06 UTC **What it missed** 605 of 2,372 items, when it started again with no help

Not measured yet

** [An agent with the slice, against an agent alone](https://featkpr.com/why/agents#measure) ** whether tests stay good as flows grow ** [Tests keeping pace with new features](https://featkpr.com/why/teams#measure) ** and the time from a change to its verdict ** [How fast a person answers](https://featkpr.com/why/teams#measure) ** one question, one panel, a default in force

## Sources

each read on 28 Sep 2026; the date is when it was published, the mark says how it was checked

- Kwa, West, Becker et al. (METR), Measuring AI Ability to Complete Long Software Tasks, NeurIPS 2025. The two figures as METR's post of 19 Mar 2025 states them; the doubling time, about every 7 months, is the paper's. 19 Mar 2025 · peer-reviewed [arxiv.org/abs/2503.14499](https://arxiv.org/abs/2503.14499) · [metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)
- Deng, Da, Pan et al. (Scale AI), SWE-Bench Pro, arXiv 2509.16941, Sep 2025. Written by the benchmark's own vendor; the paper's framing of Opus 4.1 and GPT-5. Sep 2025 · preprint, vendor [arxiv.org/abs/2509.16941](https://arxiv.org/abs/2509.16941)
- GitLab, Global DevSecOps Survey 2020 (about 3,700 respondents in 21 countries): "Last year 49% said test was at fault; this year it was 47%." The 2021 post says testing was the top cause three years running, with no percentage. GitLab sells a DevOps platform with testing built in. 18 May 2020 · vendor survey [about.gitlab.com/blog/devsecops-survey-released/](https://about.gitlab.com/blog/devsecops-survey-released/) · [about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/](https://about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/)
- P. Straubinger, G. Fraser (University of Passau), A Survey on What Developers Think About Testing, arXiv 2309.01154 (Sep 2023), ISSRE 2023 (venue checked on Crossref, 28 Sep 2026). 284 developers, recruited online; about a third were students who also worked. Sep 2023 · peer-reviewed [arxiv.org/abs/2309.01154](https://arxiv.org/abs/2309.01154) · [doi.org/10.1109/issre59848.2023.00075](https://doi.org/10.1109/issre59848.2023.00075)
- J. Micco, Flaky Tests at Google and How We Mitigate Them, Google Testing Blog: "about 84% of the transitions we observe from pass to fail involve a flaky test"; "almost 16% of our tests have some level of flakiness". 27 May 2016 · engineering blog, Google's own data [testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html)
- GitHub, Open Source Survey 2017: about 5,500 respondents from 3,800+ repositories and 500 from elsewhere. 2017 · survey [opensourcesurvey.org/2017/](https://opensourcesurvey.org/2017/)
- Alian, Nashid, Shahbandeh, Shabani, Mesbah (UBC), Feature-Driven End-To-End Test Generation (AutoE2E and E2EBench), arXiv 2408.01894, 2024. Aug 2024 · preprint [arxiv.org/abs/2408.01894](https://arxiv.org/abs/2408.01894)

## See the map on BookStack

A call of about 30 minutes: the map, a flow waiting on a verdict, and what sits on top. We read each request and write back. [Maintain an open-source app? Tell us](https://featkpr.com/demo?want=open-library) .

[Ask for a private demo](https://featkpr.com/demo)

---

The page this twin stands for: https://featkpr.com/why. Every page on this site has a `.md` twin, and answers `Accept: text/markdown`.
