Teams usually don't like testing. It slows releases and rarely catches what matters.
Problem 2 of 2, in Tim's words: developers, owners and even QA dislike tests. Suites rarely keep up with new features, what they catch is often basic, and the fix holds the release. Below: what outside research supports, what it doesn't, and what we still want to know.
read on 28 Sep 2026 · every outside figure carries its source, its date and whether it was peer-reviewed · BookStack's numbers are featkpr's records of 28 Sep 2026
featkpr holds itnot measured yet, or being builtplannedan idea, not decided
Problem 2 of 2 · Problem 1: coding agents and growing tasks · Why featkpr, the short version
The evidence #
Eight findings, ranked by how closely they match the claim. Surveys run by companies that sell testing are marked vendor.
- 47% of about 3,700 people in GitLab's 2020 survey named testing the top cause of release delays; 49% said so in 2019. GitLab reports testing first three years running. GitLab survey, 2020 · 18 May 2020 · vendor survey
- 84% of the changes from pass to fail seen at Google involved a flaky test, one that passes and fails on the same code. Almost 16% of Google's tests showed some flakiness. Micco, Google, 2016 · 27 May 2016 · engineering blog, Google's own data
- 58% of 284 developers prefer writing code to writing tests. Many would rather test less and call it boring; they would test more if managers and peers recognised it. Straubinger and Fraser, ISSRE 2023 · Sep 2023 · peer-reviewed
- 73% to 92% of browser tests needed repair when six open-source web apps moved to a newer release. The releases were months to years apart. Leotta et al., WCRE 2013 · 2013 · peer-reviewed
- 51% of 335 developers and testers meet flaky tests at least weekly, and 66% call them a moderate or serious problem. Lost trust and wasted time are the worst effects, they say. Gruber and Fraser, ICST 2022 · Mar 2022 · peer-reviewed
- 1.23% of Google's test targets ever found a real breakage in the period studied. Of 5.5 million tests that changes affected, about 63K ever failed. Memon et al., Google, 2017 · May 2017 · peer-reviewed
- half of 2,443 developers, watched in their editors for 2.5 years, ran no tests there. They spent a quarter of their time on tests, and believed it was half. Beller et al., IEEE TSE 2017 · 2017 · peer-reviewed
- broken When another team owns the tests, "test suites are frequently in a broken state", and the build stays broken until that team fixes them. DORA, Google Cloud · undated page · research-backed guidance, not a survey figure
Browser tests that needed repair after a new release, per app #
One study, one unit: the share of each app's browser tests that had to be fixed before they ran again on the newer release. Two ways of writing the same tests: written as code with WebDriver, and recorded in Selenium IDE.
written as code (WebDriver)recorded (Selenium IDE)ran again as it was
- MantisBT bug tracker · 41 tests
- PPMA password manager · 23 tests
- Claroline learning platform · 40 tests
- Address Book contacts · 28 tests
- MRBS room booking · 24 tests
- Collabtive project management · 40 tests
- All six 196 tests
Our own count: screens against end-to-end spec files in open-source apps #
How many screens each app has, and how many end-to-end spec files its maintainers committed, counted from each repository's files. 8 of 30 have no end-to-end spec file.
- OpenProject406 screens · 787 spec files
- GLPI315 screens · 107 spec files
- Redmine223 screens · 28 spec files
- Mastodon214 screens · 107 spec files
- Supabase Studio190 screens · 29 spec files
- PostHog169 screens · 43 spec files
- Discourse166 screens · 295 spec files
- Wagtail165 screens · 8 spec files
- Snipe-IT154 screens · no end-to-end spec file
- Akaunting121 screens · no end-to-end spec file
- Weblate120 screens · 1 spec file
- Gitea111 screens · 30 spec files
The other 18 apps
- Kanboard102 screens · no end-to-end spec file
- Keycloak94 screens · 81 spec files
- Grafana90 screens · 213 spec files
- Directus73 screens · 99 spec files
- Taiga72 screens · 47 spec files
- Mealie66 screens · 2 spec files
- NocoDB66 screens · no end-to-end spec file
- FreshRSS63 screens · no end-to-end spec file
- n8n61 screens · 284 spec files
- Vikunja59 screens · 72 spec files
- Umami59 screens · 8 spec files
- Kimai48 screens · no end-to-end spec file
- Linkwarden36 screens · 1 spec file
- Outline32 screens · no end-to-end spec file
- Hoppscotch32 screens · no end-to-end spec file
- Appsmith21 screens · 667 spec files
- Paperless-ngx21 screens · 5 spec files
- Uptime Kuma15 screens · 7 spec files
What the evidence does not show #
- No survey of owners. We found none that asks product owners or managers what they think of testing. That part of the claim is Tim's view.
- QA may not dislike testing. QA surveys show people stretched: lack of time was the top obstacle for 48% in Katalon's 2024 report, and changing requirements for 34%. Katalon, 2024 · vendor survey
- One QA survey points the other way. Katalon's 2025 report is framed around QA finding more joy in the work. It is a vendor survey, as is the one above. Katalon, 2025 · vendor survey
- Nothing shows that what tests catch is basic. Google's figures show how rarely tests fire, not how serious the breakage was. That part is Tim's experience. Memon et al., Google, 2017 · peer-reviewed
- Tests falling behind new features is measured only from the side. Browser tests break when a release changes the app, changing requirements are a top obstacle for QA, and code changes leave suites broken when another team owns them. We found no study that measures the lag itself.
- Late fixes are not proven to cost more. The old rule that a defect found late costs far more to fix is disputed: a study of 171 projects found no evidence for it. So we say only that a late fix holds the release, which GitLab's respondents say. Menzies et al., EMSE 2017 · peer-reviewed
- Each sample has limits. GitLab sells a platform with testing built in. About a third of the 284 developers were students who also worked, recruited online. The browser-test releases were months to years apart. Google's flaky-test figures are from 2016.
The questions we still have #
- Do owners and product managers see testing as what holds a release? No survey we found asks them.
- What do suites catch before a release: slips anyone would spot, or real regressions? How often each?
- How far behind new features does a suite fall, in days or in releases?
- Do QA people dislike testing, or the time pressure and changing requirements around it?
- When a suite catches something late, how long does the release wait?
- Would a team trust tests drafted from a map it approved more than tests written by hand?
How the map answers it: our idea, to be measured #
The team approves the map first: the features, how they lead into each other, the flows people take and the goals they reach. Each piece is one question with a default in force, so a person can keep or drop it fast. Tests, and anything else, are drafted from the approved map. When a feature is added, the map changes first and its tests follow.
Our idea is that tests drawn from a map the team already agreed keep pace with new features, and that a red result then points at a goal the team cares about. Nobody has measured it yet.
One question, with its context in one place #
- The goal, its module, where it starts and how many steps
- Read the goal, walk its steps on real screens, see the result, then decide
- One question, in plain words
- The default in force until you answer
- What keeping and dropping each do
- The goal in words, which you can rewrite
On BookStack, 149 of 149 flows are still waiting on a person's verdict (28 Sep 2026). BookStack isn't ours, so nobody there answers them. How fast a person answers has not been measured yet.
On BookStack today: 497 tests drafted from the map, by templates; on 28 Sep 2026, 06:06 UTC, 311 ran and 0 failed. Whether they keep pace with BookStack's new features is not measured yet. The tests on top of the map, in the product tour
How we will measure it #
each row names its plan item and that item's state today; what we would count is our proposal, marked as such
- Time from a change to its verdict
The minutes from a push to broke, held or not proven, and how often a verdict holds a release. our proposal
The merge-request verdict, posted with the change's own code runbeing built now
- Tests keeping pace with new features
On featkpr's own pull requests: for each feature added, the days until the map holds it and a test covers it. our proposal
featkpr crawls and tests itselfnext, est. 3 to 9 Oct
- What the tests catch
Faults planted on purpose, and how many the tests catch, with a slip and a broken goal counted apart. our proposal
Fault injectionplanned, no week yet
- What people think
The questions above, asked of developers, owners and QA on the first team, before and after. our proposal
The first team on itplanned, no week yet
- What the map misses
Measured against a list checked by hand, aiming to find 9 in 10. on the plan as written
How much featkpr misses, measuredplanned, no week yet
Sources #
each read on 28 Sep 2026; the date is when it was published, the mark says how it was checked
- GitLab, Global DevSecOps Survey 2020 (about 3,700 respondents in 21 countries): "Last year 49% said test was at fault; this year it was 47%." The 2021 post says testing was the top cause three years running, with no percentage. GitLab sells a DevOps platform with testing built in. 18 May 2020 · vendor survey about.gitlab.com/blog/devsecops-survey-released/ · about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/
- J. Micco, Flaky Tests at Google and How We Mitigate Them, Google Testing Blog: "about 84% of the transitions we observe from pass to fail involve a flaky test"; "almost 16% of our tests have some level of flakiness". 27 May 2016 · engineering blog, Google's own data testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
- P. Straubinger, G. Fraser (University of Passau), A Survey on What Developers Think About Testing, arXiv 2309.01154 (Sep 2023), ISSRE 2023 (venue checked on Crossref, 28 Sep 2026). 284 developers, recruited online; about a third were students who also worked. Sep 2023 · peer-reviewed arxiv.org/abs/2309.01154 · doi.org/10.1109/issre59848.2023.00075
- M. Leotta, D. Clerissi, F. Ricca, P. Tonella, Capture-Replay vs. Programmable Web Testing: An Empirical Assessment during Test Case Evolution, WCRE 2013. Six open-source web apps, each moved to a newer release 8 months to 4 years 8 months later; the counts per app from its Table III, summed by us. 2013 · peer-reviewed sepl.dibris.unige.it/publications/2013-leotta-WCRE.pdf
- M. Gruber, G. Fraser, A Survey on How Test Flakiness Affects Developers and What Support They Need To Address It, arXiv 2203.00483, ICST 2022. 335 developers and testers. Mar 2022 · peer-reviewed arxiv.org/abs/2203.00483
- A. Memon, Z. Gao, B. Nguyen, S. Dhanda, E. Nickell, R. Siemborski, J. Micco, Taming Google-Scale Continuous Testing, ICSE-SEIP 2017 (venue checked on Crossref, 28 Sep 2026). It measures how often tests signal a breakage, not how serious the breakage was. May 2017 · peer-reviewed static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf · doi.org/10.1109/icse-seip.2017.16
- M. Beller, G. Gousios, A. Panichella, S. Amann, S. Proksch, A. Zaidman, Developer Testing in the IDE: Patterns, Beliefs, and Behavior, IEEE TSE 2017. 2,443 engineers watched in their editors for 2.5 years. 2017 · peer-reviewed inventitech.com/assets/publications/2017_beller_gousios_panichella_amann_proksch_zaidman_developer_testing_in_the_ide_patterns_beliefs_and_behavior.pdf
- DORA (Google Cloud), Capabilities: Test automation: "Test suites are frequently in a broken state" when another team owns them. undated page · research-backed guidance, not a survey figure dora.dev/capabilities/test-automation/
- Katalon, State of Software Quality Report 2024: top obstacles for QA teams, 2022 to 2024. Katalon sells test automation. 2024 · vendor survey katalon.com/reports/state-quality-2024
- Katalon, State of Software Quality Report 2025 (1,500+ QA professionals), framed around QA finding more joy in the work. Katalon sells test automation. 2025 · vendor survey katalon.com/reports/state-quality-2025
- T. Menzies, W. Nichols, F. Shull, L. Layman, Are Delayed Issues Harder to Resolve? Revisiting Cost-to-Fix of Defects throughout the Lifecycle, Empirical Software Engineering 2017, arXiv 1609.04886: 171 projects, "no evidence for the delayed issue effect". 2017 · peer-reviewed arxiv.org/abs/1609.04886
- Bacchelli and Bird, Expectations, Outcomes, and Challenges of Modern Code Review, ICSE 2013: "code and change understanding is the key aspect of code reviewing". May 2013 · peer-reviewed www.microsoft.com/en-us/research/publication/expectations-outcomes-and-challenges-of-modern-code-review/
See the map on BookStack #
A call of about 30 minutes: the map, a flow waiting on a verdict, and the tests drafted from it. Problem 1: coding agents and growing tasks.