# Teams usually don't like testing. It slows releases and rarely catches what matters.

Problem 2 of 2, in Tim's words: developers, owners and even QA dislike tests. Suites rarely keep up with new features, what they catch is often basic, and the fix holds the release. Below: what outside research supports, what it doesn't, and what we still want to know.

read on 28 Sep 2026 · every outside figure carries its source, its date and whether it was peer-reviewed · BookStack's numbers are featkpr's records of 28 Sep 2026

featkpr holds it not measured yet, or being built planned an idea, not decided

Problem 2 of 2 · [Problem 1: coding agents and growing tasks](https://featkpr.com/why/agents) · [Why featkpr, the short version](https://featkpr.com/why)

## The evidence

Eight findings, ranked by how closely they match the claim. Surveys run by companies that sell testing are marked vendor.

- **47%** of about 3,700 people in GitLab's 2020 survey named testing the top cause of release delays; 49% said so in 2019. GitLab reports testing first three years running. [GitLab survey, 2020](https://about.gitlab.com/blog/devsecops-survey-released/) · 18 May 2020 · vendor survey
- **84%** of the changes from pass to fail seen at Google involved a flaky test, one that passes and fails on the same code. Almost 16% of Google's tests showed some flakiness. [Micco, Google, 2016](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html) · 27 May 2016 · engineering blog, Google's own data
- **58%** of 284 developers prefer writing code to writing tests. Many would rather test less and call it boring; they would test more if managers and peers recognised it. [Straubinger and Fraser, ISSRE 2023](https://arxiv.org/abs/2309.01154) · Sep 2023 · peer-reviewed
- **73% to 92%** of browser tests needed repair when six open-source web apps moved to a newer release. The releases were months to years apart. [Leotta et al., WCRE 2013](https://sepl.dibris.unige.it/publications/2013-leotta-WCRE.pdf) · 2013 · peer-reviewed
- **51%** of 335 developers and testers meet flaky tests at least weekly, and 66% call them a moderate or serious problem. Lost trust and wasted time are the worst effects, they say. [Gruber and Fraser, ICST 2022](https://arxiv.org/abs/2203.00483) · Mar 2022 · peer-reviewed
- **1.23%** of Google's test targets ever found a real breakage in the period studied. Of 5.5 million tests that changes affected, about 63K ever failed. [Memon et al., Google, 2017](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf) · May 2017 · peer-reviewed
- **half** of 2,443 developers, watched in their editors for 2.5 years, ran no tests there. They spent a quarter of their time on tests, and believed it was half. [Beller et al., IEEE TSE 2017](https://inventitech.com/assets/publications/2017_beller_gousios_panichella_amann_proksch_zaidman_developer_testing_in_the_ide_patterns_beliefs_and_behavior.pdf) · 2017 · peer-reviewed
- **broken** When another team owns the tests, "test suites are frequently in a broken state", and the build stays broken until that team fixes them. [DORA, Google Cloud](https://dora.dev/capabilities/test-automation/) · undated page · research-backed guidance, not a survey figure

### Browser tests that needed repair after a new release, per app

One study, one unit: the share of each app's browser tests that had to be fixed before they ran again on the newer release. Two ways of writing the same tests: written as code with WebDriver, and recorded in Selenium IDE.

written as code (WebDriver) recorded (Selenium IDE) ran again as it was

- **MantisBT** bug tracker · 41 tests 32 of 41 33 of 41
- **PPMA** password manager · 23 tests 17 of 23 23 of 23
- **Claroline** learning platform · 40 tests 20 of 40 40 of 40
- **Address Book** contacts · 28 tests 28 of 28 28 of 28
- **MRBS** room booking · 24 tests 23 of 24 24 of 24
- **Collabtive** project management · 40 tests 23 of 40 32 of 40
- **All six** 196 tests **73%** 143 of 196 **92%** 180 of 196

The releases were 8 months to 4 years 8 months apart, so this is a large jump, not one sprint. [Leotta et al., WCRE 2013](https://sepl.dibris.unige.it/publications/2013-leotta-WCRE.pdf) · 2013 · peer-reviewed

### Our own count: screens against end-to-end spec files in open-source apps

How many screens each app has, and how many end-to-end spec files its maintainers committed, counted from each repository's files. 8 of 30 have no end-to-end spec file.

- **OpenProject** 406 screens · 787 spec files
- **GLPI** 315 screens · 107 spec files
- **Redmine** 223 screens · 28 spec files
- **Mastodon** 214 screens · 107 spec files
- **Supabase Studio** 190 screens · 29 spec files
- **PostHog** 169 screens · 43 spec files
- **Discourse** 166 screens · 295 spec files
- **Wagtail** 165 screens · 8 spec files
- **Snipe-IT** 154 screens · no end-to-end spec file
- **Akaunting** 121 screens · no end-to-end spec file
- **Weblate** 120 screens · 1 spec file
- **Gitea** 111 screens · 30 spec files

The other 18 apps

- **Kanboard** 102 screens · no end-to-end spec file
- **Keycloak** 94 screens · 81 spec files
- **Grafana** 90 screens · 213 spec files
- **Directus** 73 screens · 99 spec files
- **Taiga** 72 screens · 47 spec files
- **Mealie** 66 screens · 2 spec files
- **NocoDB** 66 screens · no end-to-end spec file
- **FreshRSS** 63 screens · no end-to-end spec file
- **n8n** 61 screens · 284 spec files
- **Vikunja** 59 screens · 72 spec files
- **Umami** 59 screens · 8 spec files
- **Kimai** 48 screens · no end-to-end spec file
- **Linkwarden** 36 screens · 1 spec file
- **Outline** 32 screens · no end-to-end spec file
- **Hoppscotch** 32 screens · no end-to-end spec file
- **Appsmith** 21 screens · 667 spec files
- **Paperless-ngx** 21 screens · 5 spec files
- **Uptime Kuma** 15 screens · 7 spec files

counted 28 Sep 2026 · tools/library_count.py, T2's protocol made mechanical; each repository shallow-cloned at its default branch on 28 Sep 2026 · [all the counts, in the library](https://featkpr.com/library#why-h) · files, not tests

## What the evidence does not show

- **No survey of owners.** We found none that asks product owners or managers what they think of testing. That part of the claim is Tim's view.
- **QA may not dislike testing.** QA surveys show people stretched: lack of time was the top obstacle for 48% in Katalon's 2024 report, and changing requirements for 34%. [Katalon, 2024](https://katalon.com/reports/state-quality-2024) · vendor survey
- **One QA survey points the other way.** Katalon's 2025 report is framed around QA finding more joy in the work. It is a vendor survey, as is the one above. [Katalon, 2025](https://katalon.com/reports/state-quality-2025) · vendor survey
- **Nothing shows that what tests catch is basic.** Google's figures show how rarely tests fire, not how serious the breakage was. That part is Tim's experience. [Memon et al., Google, 2017](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf) · peer-reviewed
- **Tests falling behind new features is measured only from the side.** Browser tests break when a release changes the app, changing requirements are a top obstacle for QA, and code changes leave suites broken when another team owns them. We found no study that measures the lag itself.
- **Late fixes are not proven to cost more.** The old rule that a defect found late costs far more to fix is disputed: a study of 171 projects found no evidence for it. So we say only that a late fix holds the release, which GitLab's respondents say. [Menzies et al., EMSE 2017](https://arxiv.org/abs/1609.04886) · peer-reviewed
- **Each sample has limits.** GitLab sells a platform with testing built in. About a third of the 284 developers were students who also worked, recruited online. The browser-test releases were months to years apart. Google's flaky-test figures are from 2016.

## The questions we still have

- Do owners and product managers see testing as what holds a release? No survey we found asks them.
- What do suites catch before a release: slips anyone would spot, or real regressions? How often each?
- How far behind new features does a suite fall, in days or in releases?
- Do QA people dislike testing, or the time pressure and changing requirements around it?
- When a suite catches something late, how long does the release wait?
- Would a team trust tests drafted from a map it approved more than tests written by hand?

## How the map answers it: our idea, to be measured

The team approves the map first: the features, how they lead into each other, the flows people take and the goals they reach. Each piece is one question with a default in force, so a person can keep or drop it fast. Tests, and anything else, are drafted from the approved map. When a feature is added, the map changes first and its tests follow.

Our idea is that tests drawn from a map the team already agreed keep pace with new features, and that a red result then points at a goal the team cares about. Nobody has measured it yet.

### One question, with its context in one place

- One question, in plain words
- The default in force until you answer
- What keeping and dropping each do
- The goal in words, which you can rewrite

On BookStack, 149 of 149 flows are still waiting on a person's verdict (28 Sep 2026). BookStack isn't ours, so nobody there answers them. How fast a person answers has not been measured yet.

[Bacchelli and Bird, Expectations, Outcomes, and Challenges of Modern Code Review, ICSE 2013: "code and change understanding is the key aspect of code reviewing".](https://www.microsoft.com/en-us/research/publication/expectations-outcomes-and-challenges-of-modern-code-review/) · peer-reviewed

On BookStack today: 497 tests drafted from the map, by templates; on 28 Sep 2026, 06:06 UTC, 311 ran and 0 failed. Whether they keep pace with BookStack's new features is not measured yet. [The tests on top of the map, in the product tour](https://featkpr.com/product/tests)

## How we will measure it

each row names its plan item and that item's state today; what we would count is our proposal, marked as such

**Time from a change to its verdict**

The minutes from a push to broke, held or not proven, and how often a verdict holds a release. our proposal

[The merge-request verdict, posted with the change's own code run](https://featkpr.com/roadmap#merge-request-verdict) being built now

**Tests keeping pace with new features**

On featkpr's own pull requests: for each feature added, the days until the map holds it and a test covers it. our proposal

[featkpr crawls and tests itself](https://featkpr.com/roadmap#featkpr-on-featkpr) next, est. 3 to 9 Oct

**What the tests catch**

Faults planted on purpose, and how many the tests catch, with a slip and a broken goal counted apart. our proposal

[Fault injection](https://featkpr.com/roadmap#fault-injection) planned, no week yet

**What people think**

The questions above, asked of developers, owners and QA on the first team, before and after. our proposal

[The first team on it](https://featkpr.com/roadmap#first-team) planned, no week yet

**What the map misses**

Measured against a list checked by hand, aiming to find 9 in 10. on the plan as written

[How much featkpr misses, measured](https://featkpr.com/roadmap#misses-measured) planned, no week yet

## Sources

each read on 28 Sep 2026; the date is when it was published, the mark says how it was checked

- GitLab, Global DevSecOps Survey 2020 (about 3,700 respondents in 21 countries): "Last year 49% said test was at fault; this year it was 47%." The 2021 post says testing was the top cause three years running, with no percentage. GitLab sells a DevOps platform with testing built in. 18 May 2020 · vendor survey [about.gitlab.com/blog/devsecops-survey-released/](https://about.gitlab.com/blog/devsecops-survey-released/) · [about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/](https://about.gitlab.com/blog/the-software-testing-life-cycle-in-2021-a-more-upbeat-outlook/)
- J. Micco, Flaky Tests at Google and How We Mitigate Them, Google Testing Blog: "about 84% of the transitions we observe from pass to fail involve a flaky test"; "almost 16% of our tests have some level of flakiness". 27 May 2016 · engineering blog, Google's own data [testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html)
- P. Straubinger, G. Fraser (University of Passau), A Survey on What Developers Think About Testing, arXiv 2309.01154 (Sep 2023), ISSRE 2023 (venue checked on Crossref, 28 Sep 2026). 284 developers, recruited online; about a third were students who also worked. Sep 2023 · peer-reviewed [arxiv.org/abs/2309.01154](https://arxiv.org/abs/2309.01154) · [doi.org/10.1109/issre59848.2023.00075](https://doi.org/10.1109/issre59848.2023.00075)
- M. Leotta, D. Clerissi, F. Ricca, P. Tonella, Capture-Replay vs. Programmable Web Testing: An Empirical Assessment during Test Case Evolution, WCRE 2013. Six open-source web apps, each moved to a newer release 8 months to 4 years 8 months later; the counts per app from its Table III, summed by us. 2013 · peer-reviewed [sepl.dibris.unige.it/publications/2013-leotta-WCRE.pdf](https://sepl.dibris.unige.it/publications/2013-leotta-WCRE.pdf)
- M. Gruber, G. Fraser, A Survey on How Test Flakiness Affects Developers and What Support They Need To Address It, arXiv 2203.00483, ICST 2022. 335 developers and testers. Mar 2022 · peer-reviewed [arxiv.org/abs/2203.00483](https://arxiv.org/abs/2203.00483)
- A. Memon, Z. Gao, B. Nguyen, S. Dhanda, E. Nickell, R. Siemborski, J. Micco, Taming Google-Scale Continuous Testing, ICSE-SEIP 2017 (venue checked on Crossref, 28 Sep 2026). It measures how often tests signal a breakage, not how serious the breakage was. May 2017 · peer-reviewed [static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf) · [doi.org/10.1109/icse-seip.2017.16](https://doi.org/10.1109/icse-seip.2017.16)
- M. Beller, G. Gousios, A. Panichella, S. Amann, S. Proksch, A. Zaidman, Developer Testing in the IDE: Patterns, Beliefs, and Behavior, IEEE TSE 2017. 2,443 engineers watched in their editors for 2.5 years. 2017 · peer-reviewed [inventitech.com/assets/publications/2017_beller_gousios_panichella_amann_proksch_zaidman_developer_testing_in_the_ide_patterns_beliefs_and_behavior.pdf](https://inventitech.com/assets/publications/2017_beller_gousios_panichella_amann_proksch_zaidman_developer_testing_in_the_ide_patterns_beliefs_and_behavior.pdf)
- DORA (Google Cloud), Capabilities: Test automation: "Test suites are frequently in a broken state" when another team owns them. undated page · research-backed guidance, not a survey figure [dora.dev/capabilities/test-automation/](https://dora.dev/capabilities/test-automation/)
- Katalon, State of Software Quality Report 2024: top obstacles for QA teams, 2022 to 2024. Katalon sells test automation. 2024 · vendor survey [katalon.com/reports/state-quality-2024](https://katalon.com/reports/state-quality-2024)
- Katalon, State of Software Quality Report 2025 (1,500+ QA professionals), framed around QA finding more joy in the work. Katalon sells test automation. 2025 · vendor survey [katalon.com/reports/state-quality-2025](https://katalon.com/reports/state-quality-2025)
- T. Menzies, W. Nichols, F. Shull, L. Layman, Are Delayed Issues Harder to Resolve? Revisiting Cost-to-Fix of Defects throughout the Lifecycle, Empirical Software Engineering 2017, arXiv 1609.04886: 171 projects, "no evidence for the delayed issue effect". 2017 · peer-reviewed [arxiv.org/abs/1609.04886](https://arxiv.org/abs/1609.04886)
- Bacchelli and Bird, Expectations, Outcomes, and Challenges of Modern Code Review, ICSE 2013: "code and change understanding is the key aspect of code reviewing". May 2013 · peer-reviewed [www.microsoft.com/en-us/research/publication/expectations-outcomes-and-challenges-of-modern-code-review/](https://www.microsoft.com/en-us/research/publication/expectations-outcomes-and-challenges-of-modern-code-review/)

## See the map on BookStack

A call of about 30 minutes: the map, a flow waiting on a verdict, and the tests drafted from it. [Problem 1: coding agents and growing tasks](https://featkpr.com/why/agents) .

[Ask for a private demo](https://featkpr.com/demo)

---

The page this twin stands for: https://featkpr.com/why/teams. Every page on this site has a `.md` twin, and answers `Accept: text/markdown`.
