Nobody holds a map of what your app does, or how its features depend on each other.
The problem has two sides. Coding agents lose quality as tasks grow, and teams usually don't like testing. featkpr reads the code and a running copy, draws the features, the flows they form and the goals they reach, and puts each one to a person to keep or drop. Tests, and anything else, sit on top of that map.
as of 28 Sep 2026 · every outside figure carries its source and whether it was peer-reviewed · featkpr's app screens run on sample data; BookStack's numbers are its own
featkpr holds itwaiting on a person, or not measured yetnot built yetan idea, not decided
Coding agents do well on small tasks, and lose quality as the tasks grow #
problem 1 of 2
Ask a coding agent for a test of a short flow and the test is usually good. Make the flow longer, across more screens, roles and features, and the tests get worse. Research on agents in general measures the same shape:
Success by how long the task takes a person
Issues resolved, by how hard the set is
None of this research measures tests directly. Our idea, to be measured: the map hands each task one slice of the app, so the task stays about the same size however large the app grows.
Teams usually don't like testing #
problem 2 of 2
Developers, owners and even QA. Suites rarely keep up with new features, what they catch is often basic, and the fix holds the release. That is Tim's view. Outside research supports part of it:
- 47% of about 3,700 people in GitLab's 2020 survey named testing the top cause of release delays (49% the year before) GitLab survey, 2020 · 18 May 2020 · vendor survey
- 58% of 284 developers prefer writing code to writing tests Straubinger and Fraser, ISSRE 2023 · Sep 2023 · peer-reviewed
- 84% of the changes from pass to fail seen at Google involved a flaky test Micco, Google, 2016 · 27 May 2016 · engineering blog, Google's own data
No survey we found asks owners, nothing shows that what tests catch is basic, and a 2025 QA survey points the other way. Our idea, to be measured: tests drafted from a map the team approved keep pace with new features.
The idea: one map between the code and the running app #
Docs describe what the app should do, and they drift. Tests check the running app, one screen at a time. featkpr keeps the layers between: the features, how they lead into each other, the flows they form and the goals they reach. Each piece goes to a person to keep or drop.
93% of the people asked had run into incomplete or outdated docs. GitHub, 2017 · survey
- The running appwhat people see and click
docs describe it · tests check it · featkpr crawls a throwaway copy
- Flows to goalsthe paths people take to reach a goal, screen by screen
- How features relatewhich feature leads to which, and the flows that use each link; the data they share is being built
featkpr's map
- Features, by modulewhat the app can do, each named
- The coderoutes, handlers, permission checks, templates
featkpr reads it
How the map answers each side: our idea, to be measured #
- Coding agents. A task gets one slice of the map, never the whole app, so it stays about the same size as the app grows. The idea in full
- Teams and testing. The team approves the map one question at a time. Tests are drafted from the approved map, so a new feature reaches the map first and its tests follow. The idea in full
What featkpr drew on BookStack, and what it missed #
BookStack is an open-source wiki we don't own. featkpr read its code and a running copy, with nobody from BookStack helping. Then it started again with no help at all and counted what it missed.
- 312 features in 13 modules, named from the codefeatkpr's records · 28 Sep 2026
- 167 screens crawled on a running copy, signed in as Adminthe crawl · 25 Sep 2026 · commit 0f5164ec
- 149 flows drafted to 88 goals, each waiting on a personthe Flows board · 28 Sep 2026
- not held yet how features lead into each other: the Flows tree, each link with the flows that use itthe real app's Flows tree · 28 Sep 2026
- not held yet the data features sharebeing built
605 of 2,372 items missed, by reason
- no detector reads it yet295
- cut by a limit on the run199
- the step that finds it isn't built67
- needed a record that wasn't seeded28
- needed a role the crawl didn't use10
- no page links to it3
- stopped by a safety guard3
BookStack · the fresh-start audit of 26 Sep 2026 · commit 0f5164ec · 592 of them featkpr's own gaps · BookStack in the library
Map first, then anything on top #
The idea since 8 Sep: read the app once into a map, have people check it, and build everything else from the checked map. What sits on top today is the start; the slots above it come from the plan.
- Tests per route and role, each sent again as the wrong user works today
- What a change reaches, and a verdict on it being built
- What a test changed, and OpenTelemetry on runs being built
- Every state change traced, and your own telemetry from 19 Oct
- Flows run as browser tests decided, no week
- Tests in TypeScript decided, no week
- Selenium and Cucumber outputs decided, no week
- The flags and errors you already track decided, no week
- Performance budgets per goal an idea
- Load tests decided, no week
- The map, served to coding agents an idea
- A guide to your product an idea
Your review: keep, drop or rewrite each piece, one question at a time
The map: modules, features, how they lead into each other, the flows they form and the goals they reach
The data features sharebeing built
What featkpr reads: your code and a test copy that says how closely it copies real hosting
What other tools map, and who decides #
AI test tools map an app so they can write tests, and mostly decide what goes in themselves. The nearest, Momentic's app graph, lets you approve each addition; its docs say it is in alpha. Code maps read the code, and map functions and files. Analytics draw flows from real traffic, so a path nobody took isn't there. Security testers check who may reach what, API by API, without a map of features. 9354456
keeps: features or flowsdecides: a person approves each piece
AI test tools: Momentic (alpha)
keeps: features, flows and goals, linkeddecides: a person approves each piece
featkpr
keeps: tests or recorded sessionsdecides: people steer or edit it
AI test tools: Playwright
keeps: code, pages, APIs or eventsdecides: people steer or edit it
Code maps: Swimm
keeps: features or flowsdecides: people steer or edit it
3 AI test tools
Analytics: Fullstory
Research: ScenGen
keeps: tests or recorded sessionsdecides: the tool decides
3 AI test tools
keeps: code, pages, APIs or eventsdecides: the tool decides
AI test tools: mabl, Qodo
4 code maps
5 security testers
keeps: features or flowsdecides: the tool decides
AI test tools: Octomind
Analytics: PostHog, Heap
Research: AutoE2E and E2EBench
The closest work we found is AutoE2E (UBC, 2024, a preprint): it infers a web app's features and writes end-to-end tests from them. We found none that keeps a checked map, with relations and a person deciding each piece. AutoE2E, 2024 · preprint
Our idea, to be measured #
What BookStack shows today, and what nobody has measured yet. Each problem page says how we plan to measure it.
Shown on BookStack BookStack · featkpr's records of 28 Sep 2026
Not measured yet
Sources #
each read on 28 Sep 2026; the date is when it was published, the mark says how it was checked
- Kwa, West, Becker et al. (METR), Measuring AI Ability to Complete Long Software Tasks, NeurIPS 2025. The two figures as METR's post of 19 Mar 2025 states them; the doubling time, about every 7 months, is the paper's. 19 Mar 2025 · peer-reviewed arxiv.org/abs/2503.14499 · metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- Deng, Da, Pan et al. (Scale AI), SWE-Bench Pro, arXiv 2509.16941, Sep 2025. Written by the benchmark's own vendor; the paper's framing of Opus 4.1 and GPT-5. Sep 2025 · preprint, vendor arxiv.org/abs/2509.16941
- GitLab, Global DevSecOps Survey 2020 (about 3,700 respondents in 21 countries): "Last year 49% said test was at fault; this year it was 47%." The 2021 post says testing was the top cause three years running, with no percentage. GitLab sells a DevOps platform with testing built in. 18 May 2020 · vendor survey about.gitlab.com/blog/devsecops-survey-released/ · about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/
- P. Straubinger, G. Fraser (University of Passau), A Survey on What Developers Think About Testing, arXiv 2309.01154 (Sep 2023), ISSRE 2023 (venue checked on Crossref, 28 Sep 2026). 284 developers, recruited online; about a third were students who also worked. Sep 2023 · peer-reviewed arxiv.org/abs/2309.01154 · doi.org/10.1109/issre59848.2023.00075
- J. Micco, Flaky Tests at Google and How We Mitigate Them, Google Testing Blog: "about 84% of the transitions we observe from pass to fail involve a flaky test"; "almost 16% of our tests have some level of flakiness". 27 May 2016 · engineering blog, Google's own data testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
- GitHub, Open Source Survey 2017: about 5,500 respondents from 3,800+ repositories and 500 from elsewhere. 2017 · survey opensourcesurvey.org/2017/
- Alian, Nashid, Shahbandeh, Shabani, Mesbah (UBC), Feature-Driven End-To-End Test Generation (AutoE2E and E2EBench), arXiv 2408.01894, 2024. Aug 2024 · preprint arxiv.org/abs/2408.01894
See the map on BookStack #
A call of about 30 minutes: the map, a flow waiting on a verdict, and what sits on top. We read each request and write back. Maintain an open-source app? Tell us.