featkpr

Teams usually don't like testing. It slows releases and rarely catches what matters.

Problem 2 of 2, in Tim's words: developers, owners and even QA dislike tests. Suites rarely keep up with new features, what they catch is often basic, and the fix holds the release. Below: what outside research supports, what it doesn't, and what we still want to know.

featkpr holds itnot measured yet, or being builtplannedan idea, not decided

Problem 2 of 2 · Problem 1: coding agents and growing tasks · Why featkpr, the short version

The evidence #

Eight findings, ranked by how closely they match the claim. Surveys run by companies that sell testing are marked vendor.

  1. 47% of about 3,700 people in GitLab's 2020 survey named testing the top cause of release delays; 49% said so in 2019. GitLab reports testing first three years running. GitLab survey, 2020 · 18 May 2020 · vendor survey
  2. 84% of the changes from pass to fail seen at Google involved a flaky test, one that passes and fails on the same code. Almost 16% of Google's tests showed some flakiness. Micco, Google, 2016 · 27 May 2016 · engineering blog, Google's own data
  3. 58% of 284 developers prefer writing code to writing tests. Many would rather test less and call it boring; they would test more if managers and peers recognised it. Straubinger and Fraser, ISSRE 2023 · Sep 2023 · peer-reviewed
  4. 73% to 92% of browser tests needed repair when six open-source web apps moved to a newer release. The releases were months to years apart. Leotta et al., WCRE 2013 · 2013 · peer-reviewed
  5. 51% of 335 developers and testers meet flaky tests at least weekly, and 66% call them a moderate or serious problem. Lost trust and wasted time are the worst effects, they say. Gruber and Fraser, ICST 2022 · Mar 2022 · peer-reviewed
  6. 1.23% of Google's test targets ever found a real breakage in the period studied. Of 5.5 million tests that changes affected, about 63K ever failed. Memon et al., Google, 2017 · May 2017 · peer-reviewed
  7. half of 2,443 developers, watched in their editors for 2.5 years, ran no tests there. They spent a quarter of their time on tests, and believed it was half. Beller et al., IEEE TSE 2017 · 2017 · peer-reviewed
  8. broken When another team owns the tests, "test suites are frequently in a broken state", and the build stays broken until that team fixes them. DORA, Google Cloud · undated page · research-backed guidance, not a survey figure

Browser tests that needed repair after a new release, per app #

One study, one unit: the share of each app's browser tests that had to be fixed before they ran again on the newer release. Two ways of writing the same tests: written as code with WebDriver, and recorded in Selenium IDE.

written as code (WebDriver)recorded (Selenium IDE)ran again as it was

  1. MantisBT bug tracker · 41 tests 32 of 41 33 of 41
  2. PPMA password manager · 23 tests 17 of 23 23 of 23
  3. Claroline learning platform · 40 tests 20 of 40 40 of 40
  4. Address Book contacts · 28 tests 28 of 28 28 of 28
  5. MRBS room booking · 24 tests 23 of 24 24 of 24
  6. Collabtive project management · 40 tests 23 of 40 32 of 40
  7. All six 196 tests 73% 143 of 196 92% 180 of 196

Our own count: screens against end-to-end spec files in open-source apps #

How many screens each app has, and how many end-to-end spec files its maintainers committed, counted from each repository's files. 8 of 30 have no end-to-end spec file.

  1. OpenProject406 screens · 787 spec files
  2. GLPI315 screens · 107 spec files
  3. Redmine223 screens · 28 spec files
  4. Mastodon214 screens · 107 spec files
  5. Supabase Studio190 screens · 29 spec files
  6. PostHog169 screens · 43 spec files
  7. Discourse166 screens · 295 spec files
  8. Wagtail165 screens · 8 spec files
  9. Snipe-IT154 screens · no end-to-end spec file
  10. Akaunting121 screens · no end-to-end spec file
  11. Weblate120 screens · 1 spec file
  12. Gitea111 screens · 30 spec files
The other 18 apps
  1. Kanboard102 screens · no end-to-end spec file
  2. Keycloak94 screens · 81 spec files
  3. Grafana90 screens · 213 spec files
  4. Directus73 screens · 99 spec files
  5. Taiga72 screens · 47 spec files
  6. Mealie66 screens · 2 spec files
  7. NocoDB66 screens · no end-to-end spec file
  8. FreshRSS63 screens · no end-to-end spec file
  9. n8n61 screens · 284 spec files
  10. Vikunja59 screens · 72 spec files
  11. Umami59 screens · 8 spec files
  12. Kimai48 screens · no end-to-end spec file
  13. Linkwarden36 screens · 1 spec file
  14. Outline32 screens · no end-to-end spec file
  15. Hoppscotch32 screens · no end-to-end spec file
  16. Appsmith21 screens · 667 spec files
  17. Paperless-ngx21 screens · 5 spec files
  18. Uptime Kuma15 screens · 7 spec files

What the evidence does not show #

  • No survey of owners. We found none that asks product owners or managers what they think of testing. That part of the claim is Tim's view.
  • QA may not dislike testing. QA surveys show people stretched: lack of time was the top obstacle for 48% in Katalon's 2024 report, and changing requirements for 34%. Katalon, 2024 · vendor survey
  • One QA survey points the other way. Katalon's 2025 report is framed around QA finding more joy in the work. It is a vendor survey, as is the one above. Katalon, 2025 · vendor survey
  • Nothing shows that what tests catch is basic. Google's figures show how rarely tests fire, not how serious the breakage was. That part is Tim's experience. Memon et al., Google, 2017 · peer-reviewed
  • Tests falling behind new features is measured only from the side. Browser tests break when a release changes the app, changing requirements are a top obstacle for QA, and code changes leave suites broken when another team owns them. We found no study that measures the lag itself.
  • Late fixes are not proven to cost more. The old rule that a defect found late costs far more to fix is disputed: a study of 171 projects found no evidence for it. So we say only that a late fix holds the release, which GitLab's respondents say. Menzies et al., EMSE 2017 · peer-reviewed
  • Each sample has limits. GitLab sells a platform with testing built in. About a third of the 284 developers were students who also worked, recruited online. The browser-test releases were months to years apart. Google's flaky-test figures are from 2016.

The questions we still have #

  1. Do owners and product managers see testing as what holds a release? No survey we found asks them.
  2. What do suites catch before a release: slips anyone would spot, or real regressions? How often each?
  3. How far behind new features does a suite fall, in days or in releases?
  4. Do QA people dislike testing, or the time pressure and changing requirements around it?
  5. When a suite catches something late, how long does the release wait?
  6. Would a team trust tests drafted from a map it approved more than tests written by hand?

How the map answers it: our idea, to be measured #

The team approves the map first: the features, how they lead into each other, the flows people take and the goals they reach. Each piece is one question with a default in force, so a person can keep or drop it fast. Tests, and anything else, are drafted from the approved map. When a feature is added, the map changes first and its tests follow.

Our idea is that tests drawn from a map the team already agreed keep pace with new features, and that a red result then points at a goal the team cares about. Nobody has measured it yet.

One question, with its context in one place #

Flows, the review panel · featkpr on sample data
  1. The goal, its module, where it starts and how many steps
  2. Read the goal, walk its steps on real screens, see the result, then decide
  3. One question, in plain words
  4. The default in force until you answer
  5. What keeping and dropping each do
  6. The goal in words, which you can rewrite

On BookStack, 149 of 149 flows are still waiting on a person's verdict (28 Sep 2026). BookStack isn't ours, so nobody there answers them. How fast a person answers has not been measured yet.

Bacchelli and Bird, Expectations, Outcomes, and Challenges of Modern Code Review, ICSE 2013: "code and change understanding is the key aspect of code reviewing". · peer-reviewed

On BookStack today: 497 tests drafted from the map, by templates; on 28 Sep 2026, 06:06 UTC, 311 ran and 0 failed. Whether they keep pace with BookStack's new features is not measured yet. The tests on top of the map, in the product tour

How we will measure it #

  1. Time from a change to its verdict

    The minutes from a push to broke, held or not proven, and how often a verdict holds a release. our proposal

    The merge-request verdict, posted with the change's own code runbeing built now

  2. Tests keeping pace with new features

    On featkpr's own pull requests: for each feature added, the days until the map holds it and a test covers it. our proposal

    featkpr crawls and tests itselfnext, est. 3 to 9 Oct

  3. What the tests catch

    Faults planted on purpose, and how many the tests catch, with a slip and a broken goal counted apart. our proposal

    Fault injectionplanned, no week yet

  4. What people think

    The questions above, asked of developers, owners and QA on the first team, before and after. our proposal

    The first team on itplanned, no week yet

  5. What the map misses

    Measured against a list checked by hand, aiming to find 9 in 10. on the plan as written

    How much featkpr misses, measuredplanned, no week yet

Sources #

  1. GitLab, Global DevSecOps Survey 2020 (about 3,700 respondents in 21 countries): "Last year 49% said test was at fault; this year it was 47%." The 2021 post says testing was the top cause three years running, with no percentage. GitLab sells a DevOps platform with testing built in. 18 May 2020 · vendor survey about.gitlab.com/blog/devsecops-survey-released/ · about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/
  2. J. Micco, Flaky Tests at Google and How We Mitigate Them, Google Testing Blog: "about 84% of the transitions we observe from pass to fail involve a flaky test"; "almost 16% of our tests have some level of flakiness". 27 May 2016 · engineering blog, Google's own data testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
  3. P. Straubinger, G. Fraser (University of Passau), A Survey on What Developers Think About Testing, arXiv 2309.01154 (Sep 2023), ISSRE 2023 (venue checked on Crossref, 28 Sep 2026). 284 developers, recruited online; about a third were students who also worked. Sep 2023 · peer-reviewed arxiv.org/abs/2309.01154 · doi.org/10.1109/issre59848.2023.00075
  4. M. Leotta, D. Clerissi, F. Ricca, P. Tonella, Capture-Replay vs. Programmable Web Testing: An Empirical Assessment during Test Case Evolution, WCRE 2013. Six open-source web apps, each moved to a newer release 8 months to 4 years 8 months later; the counts per app from its Table III, summed by us. 2013 · peer-reviewed sepl.dibris.unige.it/publications/2013-leotta-WCRE.pdf
  5. M. Gruber, G. Fraser, A Survey on How Test Flakiness Affects Developers and What Support They Need To Address It, arXiv 2203.00483, ICST 2022. 335 developers and testers. Mar 2022 · peer-reviewed arxiv.org/abs/2203.00483
  6. A. Memon, Z. Gao, B. Nguyen, S. Dhanda, E. Nickell, R. Siemborski, J. Micco, Taming Google-Scale Continuous Testing, ICSE-SEIP 2017 (venue checked on Crossref, 28 Sep 2026). It measures how often tests signal a breakage, not how serious the breakage was. May 2017 · peer-reviewed static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf · doi.org/10.1109/icse-seip.2017.16
  7. M. Beller, G. Gousios, A. Panichella, S. Amann, S. Proksch, A. Zaidman, Developer Testing in the IDE: Patterns, Beliefs, and Behavior, IEEE TSE 2017. 2,443 engineers watched in their editors for 2.5 years. 2017 · peer-reviewed inventitech.com/assets/publications/2017_beller_gousios_panichella_amann_proksch_zaidman_developer_testing_in_the_ide_patterns_beliefs_and_behavior.pdf
  8. DORA (Google Cloud), Capabilities: Test automation: "Test suites are frequently in a broken state" when another team owns them. undated page · research-backed guidance, not a survey figure dora.dev/capabilities/test-automation/
  9. Katalon, State of Software Quality Report 2024: top obstacles for QA teams, 2022 to 2024. Katalon sells test automation. 2024 · vendor survey katalon.com/reports/state-quality-2024
  10. Katalon, State of Software Quality Report 2025 (1,500+ QA professionals), framed around QA finding more joy in the work. Katalon sells test automation. 2025 · vendor survey katalon.com/reports/state-quality-2025
  11. T. Menzies, W. Nichols, F. Shull, L. Layman, Are Delayed Issues Harder to Resolve? Revisiting Cost-to-Fix of Defects throughout the Lifecycle, Empirical Software Engineering 2017, arXiv 1609.04886: 171 projects, "no evidence for the delayed issue effect". 2017 · peer-reviewed arxiv.org/abs/1609.04886
  12. Bacchelli and Bird, Expectations, Outcomes, and Challenges of Modern Code Review, ICSE 2013: "code and change understanding is the key aspect of code reviewing". May 2013 · peer-reviewed www.microsoft.com/en-us/research/publication/expectations-outcomes-and-challenges-of-modern-code-review/

See the map on BookStack #

A call of about 30 minutes: the map, a flow waiting on a verdict, and the tests drafted from it. Problem 1: coding agents and growing tasks.

Ask for a private demo

Open original

↑ ↓ to move, Enter to open, Esc to close