The roadmap on one time axis: three weeks shipped, five planned.
One lane per part of the product. Shipped work sits on the day it landed, this week's work at today, planned weeks after it. Select any node or bar for the same detail as its card on the board.
as of 28 Sep 2026 · 12 shipped · 3 being built · 7 in planned weeks · 20 decided, no week · 18 ideas
Dates ahead are estimates. On 27 Sep, effects and OpenTelemetry landed a week early and the merge-request verdict was pulled forward. On 28 Sep featkpr's verdict on its own pull requests and the release-flow editor landed about a week ahead of their estimates, so expect this page to change.
Select a node or a bar for its details. On a keyboard, Tab reaches each one; ← and → move along a lane.
the map2 dated, 7 not
flows2 dated, 4 not
decisions3 dated, 4 not
on top: tests and proof4 dated, 11 not
on top: changes and releases5 dated, 0 not
platform and the library6 dated, 12 not
A filled node shipped on that day · an open bar is being built this week · a dashed bar is a planned week, an estimate · the shaded part is ahead of today. Commits: commits on the main development line, per day (git log; 28 Sep counted on dev to 06:33 EDT).
Shipped 28 Sepon top: changes and releases
featkpr reads its own code and reviews its own pull requests
featkpr maps its own Python and React code and posts a verdict on its own pull request.
- Why
- A second stack shows the readers are not shaped around BookStack, and reviewing its own changes every day is the quickest way to find where featkpr is wrong.
- What it made possible
- featkpr now maps itself, a second stack (Python and React): 207 candidates and 39 goals read from its own code, with no model call. Its first verdict went up on its own pull request #798.
- What we know
- 207 candidates, 39 goals and 108 code anchors read from featkpr's own code, with no model callthe production store ·
- Its surfaces, counted in the repository: 40 API operations, about 24 screens, 72 commandsthe product-three audit ·
- The pull request's code is only read; the shipped service writes the verdict, so the tool never judges itselfthe product-three design ·
- What shipped
- 207 candidates read from featkpr's own API, commands and screens
- 39 goals drawn from its own code
- first verdict on its own pull request #798: pass, 0 goals touched
- its own changes go dev → a daily release pull request → prod, and featkpr reviews the release
- Python (FastAPI routes, click commands) and React screens read; GitHub repositories followed by webhook
- what reviewers said, from people, bots and agents, kept as evidence beside the verdict
Shipped 28 Sepon top: changes and releases
How your app ships, drawn and approved
featkpr reads how an app ships, proposes the flow, and a person approves the drawing.
- Why
- "Proven" needs a place: proven on the development branch is not proven in the release. Knowing the flow lets every verdict say which stage it speaks for, and a release's verdict cover everything merged since the last one.
- What it made possible
- featkpr reads how an app ships (branch names, tags, merge history), proposes one of nine flows with its evidence, and a person fixes the drawing and approves it. The approved drawing then shows what featkpr does on each arrow.
- What we know
- Nine flows it can propose, drawn from Kargo's stages, Sentry's releases, Octopus lifecycles and GitHub environmentsrelease-flow research ·
- BookStack has no promotion pull requests: development is merged into release, then tagged; featkpr reads it as a flow of its ownBookStack's branches and tags, read from Codeberg ·
- Three editor designs drawn on real flows; the drafting board chosen, with an outline view for keyboard usersthe release-flow design pass ·
- What shipped
- nine flows it can propose, each with its evidence
- BookStack read as development → release by a pushed merge, then tags
- a release's verdict covers everything merged since the last one
- the map follows BookStack's development branch, never a pull request's head; names carried across the move for less than a cent

Shipped 28 Sepon top: tests and proof
A test copy that behaves like real hosting
A test copy with real mail and a real queue, and a report of how close it is to hosting.
- Why
- A test that passes on a copy unlike real hosting says little about real hosting. Reporting how close each copy is, on every run, lets a goal say where it was proven.
- What it made possible
- The copy featkpr tests now sends real mail, runs a real queue, and says how closely it copies real hosting: 11 of 17 checks on the production-like tier, 4 of 17 on the everyday one, with the same results on both.
- What we know
- Production-like copy: 11 of 17 checks match real hosting, the everyday copy 4; 311 tests passed on both, and none changed result between themthe comparison at BookStack's development tip ·
- Mail of all 7 kinds caught in one run, where the day before none wasthe run at BookStack's development tip ·
- Only the copy with a real queue showed that a queued comment mail for a page deleted before it sends failsthe production-like run ·
- An independent review of the first design found four places where it would have made verdicts wrong, among them users hard-coded so the wrong-user runs could silently skipthe environments review ·
- What shipped
- mail really sent and caught: 7 of 7 kinds, a new user's sign-up included
- people as personas with a role and a mailbox; the wrong user is derived, never skipped
- a fidelity report on every run
Shipped 27 Sepdecisions
Goals that carry their proof, and the first merge-request verdict
Each goal keeps its own record of proof, and the first verdict on a real pull request.
- Why
- A goal is what a person wants to reach, such as publishing a page. When each goal carries its proof, a change is judged by the goals it touches, not by the lines it edits.
- What it made possible
- A goal now keeps its own record of proof, computed from runs. Each drafted goal is checked by a second model, and its code anchors are put to a vote of three cheap models, before a person sees it. The first verdict on a real pull request went up: BookStack's #6213, in a private mirror. It read "needs a look, 2 goals touched, 0 proven, nothing ran": on 27 Sep featkpr did not yet run a pull request's own code. It did on 28 Sep.
- What we know
- A checking model caught 46 of 48 overclaiming goal sentences across the drafts testedthe drafting-model experiment ·
- A vote of three cheap models caught 94% of planted anchor errors with 2 false alarms, where one model alone raised about 40the drafting-model experiment, on planted cases ·
- The first checked run: 103 of 168 goals walked by a flow, 243 flows in allthe register ·
- First verdict on BookStack #6213: needs a look, 2 goals touched, 0 proven, nothing ranfeatkpr's private mirror ·
- What shipped
- each goal's record of proof, live 27 Sep 14:30 UTC
- goals checked by a second model and a vote of three, 103 of 168 walked by a flow
- how complex each goal is, drafted from what drives it; a person's word is feedback on the draft
- the first merge-request verdict, on BookStack's pull request #6213
- replies to the verdict's comment are read back, shipped 27 Sep (replies arrive since 27 Sep night)
- built and shipped: what a test changed (effects) and zero-code OpenTelemetry; effects recorded on BookStack's run of 28 Sep
- the crawl may edit and delete what it made: verified on a test copy of BookStack on 27 Sep, 4 edited, 4 deleted, nothing else touched
Shipped 26 Sepon top: tests and proof
The first tests run, with the wrong-user check
Each test is sent again as the wrong user, and must then fail.
- Why
- A passing test only means something if it can fail. Sending each test again as someone who should get the opposite answer proves the permission check exists, and turns a green run into evidence.
- What it made possible
- Proof that a permission check exists: each test is sent again as the wrong user, someone who should get the opposite answer, and must then fail. And a check that lists what featkpr missed when it started again with no help.
- What we know
- 304 tests ran and passed, 0 failed; 269 of them proven by the wrong-user checkBookStack's run at 0f5164ec, 17:42 UTC ·
- Run again at BookStack's development tip: 311 tests, 276 proven, 0 failedthe production store, exported ·
- Starting again with no help, featkpr missed 605 of 2,372 items; the check lists each one with its reasonthe fresh-start audit ·
- What shipped
- every test passed on the first run, 0 failed, and most were proven by the wrong-user check (the counts are in the report of 26 Sep)
- the Runs page
- the missed-items check: 605 of 2,372 missed

Shipped 25 Sepflows
The crawl sends forms
The crawl fills in forms, and flows are drafted from what it really did.
- Why
- A flow drawn from real screens can be judged by a person in seconds: they see the path and keep it or drop it. Sending forms is what lets the crawl reach the screens behind a save.
- What it made possible
- Flows drafted from what the crawl really did: a goal, then the flow a person takes to it, screen by screen, for a person to keep or drop.
- What we know
- 167 screens captured and 8 forms sent, signed in as Adminthe admin crawl at 0f5164ec ·
- 149 flows to 88 goals now wait on a person's keep or dropthe app's Flows screen ·
- What shipped
- 167 screens captured as Admin
- flows drafted to goals

Shipped 24 Sepplatform and the library
Sign-in, and a second app
Sign-in for owners, and Shlink read and named beside BookStack.
- Why
- A pipeline built around one app proves little. A second app on another stack shows which parts are general and which were only BookStack's.
- What it made possible
- Shlink beside BookStack in one deployment: the first sign that the pipeline is not built for one app.
- What we know
- Shlink read from its code: 20 features named from 28 routesShlink's map export ·
- What shipped
- the products screen and switcher
- Shlink read and named
Shipped 22 Septhe map
The first map, and the name featkpr
BookStack's routes grouped into named features, module by module.
- Why
- The map is what every later step hangs on: flows are drawn across its features, tests are written per feature, and a change is traced to the features it reaches.
- What it made possible
- BookStack's routes, grouped into features by module, for $0.99 of model calls. The map every later step hangs on.
- What we know
- BookStack's map today: 312 features in 13 modulesthe production store, exported ·
- Of the routes BookStack's own route list names, the reader dropped 16, all console commands no reader handles yetthe fresh-start audit ·
- What shipped
- features named by module
- the name featkpr

Shipped 21 Sepdecisions
The first build goes live
The web app and its API, with the decisions waiting on a person as cards.
- Why
- A product owner needs one place that shows what waits on them. The screens were designed as one system before any was built, so every screen added since has had a place to go.
- What it made possible
- An app a product owner opens: the screens were designed as one system first, then built.
- What we know
- Three whole design candidates drawn and compared before building; the simplest correct one chosen, with the best ideas of the other two grafted inthe design choice ·
- What shipped
- the web app and its API
- decisions waiting on a person, as cards

Shipped 18 Sepplatform and the library
The specification and its test plan
A written design with three rules, and a test plan every part is checked against.
- Why
- Three rules keep the product honest: nothing a model judges changes a mechanical result, a suspected bug is never quietly accepted as the new normal, and naming organises features but never deletes one. The test plan holds every part to them before it merges.
- What it made possible
- Every part of the build is checked against a written test plan before it merges.
- What we know
- The naming pass raised precision from 0.26–0.44 to 0.764; allowed to drop candidates, it lost 21 of 130 features, so naming may never deletethe design, measured on BookStack ·
- On the one real regression in BookStack's history, a model judge called it "just a rename" three times out of three; the mechanical check got it rightthe design, measured on BookStack ·
- Nine templates wrote tests for 20 of 20 revision cases with no model call; all 20 passed, and all 20 failed as the wrong userthe design, section 9 ·
Shipped 17 Sepplatform and the library
Fourteen research reports, and a first app
Fourteen research reports, and BookStack chosen as the first app to measure.
- Why
- A claim about one app is only worth something if the next app can be counted the same way. The research fixed a counting method first, so every later number on this site has a method behind it.
- What it made possible
- BookStack chosen, and a counting protocol every later app is measured on, so a claim about one app can be checked on the next.
- What we know
- The market map found no vendor holding a map of an app's features and how they relate; its verdict was to go ahead on the mapresearch report T1, the market map ·
- BookStack counted on a ten-count protocol: 93 screens, 77 API operations, 13 modulesresearch report T2, the counting protocol ·
- A commercial AI testing agent run on BookStack as the line to beat: it reached 40 of 93 screens and 6 of 42 features, and no permission gateresearch report G1.1, a trial run ·
- What shipped
- BookStack counted at its v26.05.5 release
- a shortlist of the next three
Shipped 8 Sepplatform and the library
The project starts
One question: can a tool list what a web app does, and prove it on the running app?
- Why
- Most teams hold no written record of what their product does: documentation drifts and tests only look at the running end. The project set out to build that record first and to put everything else on top of it.
- What it made possible
- The question: can a tool list what a web app does from its code, and prove it on the running app?
- What we know
- The repository's first commit, two days ingit history of the repository ·
Now, this weekplatform and the library
Add an app: paste its address, give one token, follow the setup list
Paste an app's address, give one token, and follow a setup list beside the map.
- Why
- Today adding an app takes hand-written setup files and a deploy, so only we can do it. When it is a paste and a token, anyone can bring their own app and watch its first map fill in.
- Where it stands
- decided 28 Sep; apps and their settings become stored records (built, in review); the setup list beside the map is designed (approved 28 Sep) and being built; until it ships we set up each app ourselves
- What we know
- Adding featkpr as a product took 7 hand-written files, a database migration and a deployonboarding research ·
- A token page with its fields filled in ahead works on GitHub, GitLab and Codeberg; Bitbucket and Azure DevOps get a link and a checklistonboarding research, platform docs ·
- Reading apps it had never seen: a FastAPI template from 0 to 23 of 23 routes, a Laravel app from 0 to 191the new reader, measured ·
- Three setup shapes compared (a wizard, a checklist that ticks itself, a setup pull request); the checklist fits setup that waits on things outside featkpronboarding research ·

Now, this weekon top: changes and releases
The merge-request verdict, posted with the change's own code run
The change's own code run beside the code it replaces, and a verdict on every push.
- Why
- Before merging, a reviewer wants to know which goals a change touches and whether their tests still hold. Running the change's own code beside the code it would replace turns "looks fine" into broke, held or not proven.
- Where it stands
- the pull request's own code ran on BookStack #6213 on 28 Sep (0 broke, 19 held, 5 not proven); next, it posts on every push and fills the Goal sheets page
- What we know
- BookStack #6213: 0 broke, 19 held, 5 not proven; of those, 4 had no runnable test before the change and 1 passes even as the wrong userthe run of the change's own code ·
- An audit of the earlier pipeline: 28 verdicts written, all on requests already merged or closed; none ran a test or named a goalthe verdict audit ·
- A survey of pull-request tools found none that speaks in goals, shows a test proven able to fail, uses runtime effects as evidence and says what it cannot seea survey of vendor docs and public pull requests ·
- Five states with exact rules, and never green when nothing ranthe verdict design pass ·

I need this featkpr.com/roadmap/timeline#merge-request-verdict
Now, this weekon top: tests and proof
Effects and OpenTelemetry on BookStack's runs
What each test step changed behind the page: rows, mail, files and outside calls.
- Why
- A page can say "saved" while nothing was saved. Recording what each step changed behind the page lets a test prove the outcome, not only the screen.
- Where it stands
- effects recorded on BookStack's run of 28 Sep at its development tip, not shown here yet; OpenTelemetry measured on one capture run, not on every run yet
- What we know
- BookStack's run at its tip recorded 1,688 activity-log entries, 518 database rows, 232 mails and 100 filesthe run at BookStack's development tip ·
- OpenTelemetry measured on one capture run, with no change to BookStack's codethe build log ·
Next, 3 to 9 Oct, estimateon top: changes and releases
featkpr crawls and tests itself
A copy of featkpr walked in isolation, then its own tests on its own pull requests.
- Why
- When featkpr's own tests say broke or held on its own pull requests, every change to featkpr becomes a live test of featkpr.
- Where it stands
- a copy of featkpr walked in isolation, signed in as the owner and as nobody, est. 3–6 Oct; then its own tests say broke or held on its own pull requests, est. 6–9 Oct
- What we know
- Planned: the crawl of a throwaway copy, signed in as the owner and as nobody, est. 3–6 Oct; its own tests on its pull requests, est. 6–9 Octthe product-three plan ·
- The crawl may send forms on the throwaway copy only, never against productionthe product-three plan ·
Next, from 12 Oct, estimateon top: changes and releases
Where a goal counts as proven: a stage and an environment per product
Each goal says where it was proven: which stage, on which kind of copy.
- Why
- "Proven" alone hides where. A goal that reads "proven · production-like at prod" tells a reviewer what the proof covers, and old verdicts replay under the rule in force when they were made.
- Where it stands
- the flow is live; next, each goal says where it was proven, e.g. 'proven · production-like at prod', never a plain 'proven'; Kargo later
- What we know
- First rules proposed: BookStack at its release, on the published release image; featkpr at prod, on the production-like copythe environments design ·
- Kargo's stages and "freight" borrowed for which commit each stage runsrelease-flow research ·
I need this featkpr.com/roadmap/timeline#proven-per-release-flow
Next, from 12 Oct, estimatethe map
Goals found in the code
Each goal's inputs, outside calls and the records it changes, read from the code.
- Why
- Goals drafted from screens miss what the code knows: which record changes, what must exist first, what undoes it. Reading that from the code links goals to each other without a model.
- Where it stands
- featkpr finds each goal's inputs and outside calls in the code itself
- What we know
- BookStack's code names 27 entities, 82 relations and 55 audit types; 54 of 164 goals map to a real change of an entitygoal-model research, measured ·
- Hand-checked: the resolver right on 58 of 60, filtered "enables" links 20 of 20, undo pairs 17 of 19goal-model research, one reviewer ·
- Letting a model invent summary goals from the bottom up misses 49–64% of themUCRBench ·
Next, from 12 Oct, estimateplatform and the library
Keeps working when the code host is down
A queue, a mirror of the code host and an operations page.
- Why
- Verdicts have to arrive when the code host is slow or down. A local mirror and a queue that retries let featkpr keep reading changes and catch up by itself.
- Where it stands
- a queue, a mirror and an operations page
- What we know
- Anonymous traffic to Codeberg stalled BookStack's chain until a token was addedonboarding research ·
- A full clone per request named as a risk of the verdict, with a mirror planned for itthe verdict plan ·
I need this featkpr.com/roadmap/timeline#survives-host-outages
Next, from 19 Oct, estimateflows
Viewer and guest flows
Flows for every role, not only the admin.
- Why
- What a viewer or a guest can reach is where permission bugs hide. Flows per role show one goal as each person sees it.
- Where it stands
- the crawl signs in as the admin today; other roles once Flows filters by role
- What we know
- A crawl as 8 roles: 144 routes fall into 22 patterns, and the code's permission checks explain 46 of the 47 role splitsthe branching study, measured on BookStack ·
- The crawl signs in as Admin todaythe admin crawl ·
I need this featkpr.com/roadmap/timeline#viewer-and-guest-flows
Next, from 19 Oct, estimateon top: tests and proof
Every state change traced during a run
Every state change traced during a run, and your product's own telemetry taken in.
- Why
- Effects say what a step changed; tracing says how the whole system moved while it did. Slow queries, stuck jobs and silent errors show there.
- Where it stands
- and your product's own telemetry, any stack
- What we know
- Effects already recorded per step: rows, mail, files and calls to other hostsBookStack's run at its tip ·
I need this featkpr.com/roadmap/timeline#whole-system-tracing
Next, from 19 Oct, estimatedecisions
The goals view
A screen per goal: its proof, its flows and how complex it is.
- Why
- Owners think in goals, not routes. One screen per goal puts its proof, its flows and what it depends on in one place, so a decision takes seconds.
- Where it stands
- a screen per goal with its proof and complexity; its design comes first. For one change, the Goal sheets page already shows each goal it touches
- What we know
- Three meanings of "covered" live in the app today, which is how one card read 0 while every test had passedgoal-model research ·
- The goal checker flagged 41 goals for a person; nothing yet puts them in front of onethe overnight run ·
Later, before the first teamon top: tests and proof
A wrong user for every test
A seeded user without the permission, for each test that has none today.
- Why
- A test with no wrong user passes whoever runs it, so it proves nothing about permissions. A seeded user without the permission gives every test someone to fail as.
- In short
- a seeded user without the permission for each test that has none today, so no test stays unchecked
- What we know
- 35 of 311 tests at BookStack's tip still have no wrong userthe production store, exported ·
- Content export is held by every seeded role, so no one can fail as that user yetthe environments build ·
I need this featkpr.com/roadmap/timeline#wrong-user-every-test
Later, before the first teamon top: tests and proof
Tests through a browser
The flows a person keeps, run as Playwright tests in a real browser.
- Why
- Tests call the app directly today. Running the kept flows in a browser proves the path a person actually sees, screen by screen.
- In short
- the flows a person keeps, run as Playwright browser tests
- What we know
- Research prototypes already infer an app's features to write browser tests: AutoE2E reports 79% feature coverage on its own benchmarkAlian et al., arXiv 2408.01894 ·
Later, before the first teamthe map
Node and Django readers
Readers for Node and Django apps, starting with Uptime Kuma and Paperless-ngx.
- In short
- Uptime Kuma and Paperless-ngx first; Python (FastAPI, click) and React readers shipped 28 Sep
I need this featkpr.com/roadmap/timeline#node-python-readers
Later, before the first teamplatform and the library
InvenTree, Immich and Formbricks
InvenTree, Immich and Formbricks join the library.
- Why
- Three more apps on different stacks test whether featkpr's readers and crawl hold up beyond the two it knows.
- In short
- the next three in the library
- What we know
- InvenTree, Immich and Formbricks named as the next targets on the counting protocolresearch report T2 ·
Later, before the first teamthe map
How much featkpr misses, measured
How much featkpr misses, measured against a checked list; aiming to find 9 in 10.
- Why
- A map that never says what it misses cannot be trusted. Measuring misses against a list checked by hand turns "we found a lot" into a number with a method.
- In short
- aiming to find at least 9 in 10
- What we know
- Starting again with no help, featkpr missed 605 of 2,372 itemsthe fresh-start audit ·
- The gate set in research: find at least 9 in 10 of a 130-row list checked by handresearch report G1.3 ·
Later, before the first teamon top: tests and proof
Tests written in TypeScript too
Tests written in TypeScript beside the pytest ones.
- In short
- beside pytest
Later, before the first teamdecisions
Owners on features
Each feature with the person who answers for it.
Later, for teamsplatform and the library
Apps registered with each platform, and invites
A GitHub App and a Forgejo integration in place of tokens, with invites and roles.
- In short
- a GitHub App and a Forgejo integration in place of tokens; invites and roles
I need this featkpr.com/roadmap/timeline#platform-apps-invites
Later, for teamsplatform and the library
Teams and tenancy
Several teams on one featkpr, each seeing only its own apps.
Later, for teamsplatform and the library
The reader in your own CI
The reader runs in your own CI, so your code stays with you.
- In short
- your code stays with you
Later, for teamson top: tests and proof
Selenium and Cucumber outputs
Tests written out for Selenium and Cucumber too.
Later, for teamsplatform and the library
The first team on it
The first team using featkpr on its own app.
Later, researched onlyon top: tests and proof
Stand-ins for SSO, LDAP and webhooks
Stand-ins for single sign-on, LDAP and webhooks in the test copy.
Later, researched onlyon top: tests and proof
Fault injection
Faults injected on purpose, to see what breaks and what holds.
Later, researched onlyon top: tests and proof
Reading the flags and errors you already track
Read the feature flags and errors you already track: GrowthBook, PostHog, Sentry.
- In short
- GrowthBook, PostHog, Sentry
Later, researched onlyon top: tests and proof
Load tests with k6
Load tests written with k6 from the same map.
Later, decided 27 Sep, not datedplatform and the library
A private read-only demo
A read-only featkpr for people you invite, behind a one-time code.
- Why
- Seeing featkpr on a real product is the proof. A read-only demo lets invited people look on their own time, once the proof is worth showing.
- In short
- for people you invite, behind a one-time code, once the proof is worth showing; until then, the call is the demo
- What we know
- Demos planned with generated data only, reset when a viewer's session startsthe environments design ·
Later, decided 27 Sep, not datedthe map
Apps made of several services or repositories
Apps made of several services or repositories, linked by the calls between them.
- Why
- A goal already spans services; what breaks is "one repository at one commit". Linking a screen to the API it calls lets a change in one service reach the goals it touches in another.
- In short
- featkpr's own services first, linked by the 31 exact calls its screens make to its API; then Shlink and its separate web client; then a different version per environment
- What we know
- featkpr's screens call its API through 31 exact typed calls, so the links between them are exactthe multi-service study ·
- Shlink's server and its separate web client are a real two-repository case, already read on one sidethe multi-service study ·
Later, decided 27 Sep, not datedflows
Every way into a goal, not just the shortest
A flow for each real way into a goal, not only the shortest.
- Why
- People arrive at a goal from a link in a mail, a signed-out page or a refusal, not only from Home. Each real way in is a path a test should walk.
- In short
- a flow per distinct way in: a deep link signed in or out, a refusal with its message, the page after a save; one goal seen per role, per feature and as a matrix
- What we know
- The shortest path started through a "Recently viewed" list or the hidden profile menu in 72 of 90 cases, which a new user does not havethe branching study, measured on BookStack ·
- 11 ways in that are not links at all, such as a deep link, a signed-out resume or a token in a mailthe branching study ·
Later, decided 27 Sep, not datedon top: tests and proof
Test data from your own exports, masked
Tests run on a masked copy of your own export; pull requests and demos get generated data.
- Why
- Tests on realistic data find what tests on ten sample rows cannot. Masking with a key per product keeps records linked while no real name reaches a test.
- In short
- a real export masked with a key per product, only on trusted nightly runs; pull requests and demos get generated data
- What we know
- Masked data only on trusted nightly runs; code from outside contributors and demos never see itthe environments design ·
- The popular data synthesiser moved to a non-open licence, so generation uses plain statistics with a fixed seedthe environments design ·
Idea, raised by Tim 28 Septhe map
Examples from much larger products: Blender, LibreOffice
The map and one scenario's slice shown on much larger products.
- Why
- featkpr's idea is that each task stays about as small however large the app is. Products the size of Blender or LibreOffice would put that idea to a real test.
- The idea
- the map and a scenario's slice shown on products far bigger than BookStack, and not web apps, to test whether each task stays bounded; nothing read yet
- What we know
- Agents succeed on almost every task that takes a person under 4 minutes, and on under 10% of tasks over about 4 hoursMETR, NeurIPS 2025 ·
- Accuracy per step falls as steps pile up, even when the plan is givenSinha et al., ICLR 2026 ·
Idea, raised by Tim 27 Sepplatform and the library
An open library of every open-source app's features
The features featkpr reads from open-source apps, public, with a free program for the projects.
- Why
- Every open-source app deserves a map of what it does. A public library grows with each app featkpr reads, and gives the projects their map and tests back.
- The idea
- the features featkpr reads from open-source apps, public and growing, and a free program for the projects themselves (a grant or a partnership) to get their map and tests. Long term.
- What we know
- BookStack's 312 features and Shlink's 20 are searchable in the library todaythe library ·
Idea, raised by Tim 27 Sep; part of it runsplatform and the library
A demo environment per feature
A throwaway copy of the app per feature, with a read-only link for clients.
- Why
- A client wants to try the feature, not read about it. A copy per feature, reset for each viewer, shows it working without risk to anything real.
- The idea
- with a read-only link for clients; the throwaway copy it needs runs today
- What we know
- Throwaway copies run today: each run gets its own, labelled, and a cleanup removes leftoversthe environments build ·
I need this featkpr.com/roadmap/timeline#demo-env-per-feature
Idea, raised by Tim 27 Sep; part of it runsdecisions
Several models check every answer
Every model answer checked by other models, not only the goals.
- Why
- One model's answer is a guess; agreement between models is evidence. Checking everywhere a model answers keeps the drafts a person sees honest.
- The idea
- a second model checks every drafted goal, and three cheap models vote on its anchors; the idea is to do it everywhere a model answers
- What we know
- On anchors, a vote of three caught 94% of planted errors with 2 false alarms, where one model alone raised about 40the drafting-model experiment ·
- The weak spot: adding a missing anchor, right 10 times in 33, so the vote may drop but never addthe drafting-model experiment ·
I need this featkpr.com/roadmap/timeline#model-vote-everywhere
Idea, raised by Tim 27 Sepon top: tests and proof
Performance budgets per goal
Page speed and web vitals per goal, once browser flows exist.
- The idea
- page speed and web vitals, once browser flows exist
I need this featkpr.com/roadmap/timeline#performance-budgets
Idea, raised by Tim 27 Sepflows
Process mining on real traces
The flows people really take, mined from real traces.
- Why
- Drafted flows are what the app allows; real traces show what people do. Comparing them finds the paths nobody tests.
- The idea
- the flows people really take, once real traces exist
- What we know
- Kept as one half-day offline experiment, once real traces existthe branching design ·
Idea, raised by Tim 27 Sepflows
A phone crawl lane
The same flows walked at phone width.
- The idea
- the same flows at phone width
Idea, raised by Tim 25 Sepdecisions
A conversation that draws the flow
A conversation beside a canvas that draws the flow you describe.
- The idea
- a conversation beside a canvas that draws what you are talking about
I need this featkpr.com/roadmap/timeline#conversation-canvas
Idea, raised by Tim 22 Septhe map
A product wiki and a flags hub
What the product does, and what is switched on, in one wiki.
- The idea
- what the product does, and what is switched on
Idea, raised by Tim 22 Septhe map
Features from every source
Features read from code, docs, release notes and feedback, not code alone.
- The idea
- code, docs, release notes and feedback, not only code
I need this featkpr.com/roadmap/timeline#features-every-source
Idea, raised by Tim 22 Sep; part of it runsplatform and the library
Findings filed with maintainers
What featkpr finds in an open-source app, sent to its maintainers after a review.
- Why
- Finding a bug in someone's app is only useful if they hear about it. Filing reviewed findings upstream gives the projects something back for being read.
- The idea
- what featkpr finds in an open-source app, sent to its maintainers after a review; one candidate found 28 Sep on the production-like copy (a queued comment mail for a deleted page fails); nothing filed yet
- What we know
- One candidate found: a queued comment mail for a page deleted before it sends fails in BookStack; nothing filed yetthe production-like run ·
I need this featkpr.com/roadmap/timeline#findings-to-maintainers
Idea, raised in the design notes 22 Septhe map
A read-only server for agents
The map served read-only to coding agents over MCP.
- The idea
- the map, served to coding agents over MCP
Idea, raised by Tim 21 Sepflows
A guide to your product, from the map
A guide to your product from the map: steps with screens, who can do what.
- The idea
- steps with screens, who can do what, where it fails
Idea, raised by Tim 26 Sepdecisions
Quick questions in a corner of the screen
A decision answered in a corner of the screen you are on.
- The idea
- a decision answered where you are, instead of on its own screen
Idea, raised by Tim 27 Sepplatform and the library
A public demo
A public demo; whether it comes before or after the private one is open.
- The idea
- open question: whether a public demo comes before or after the private read-only one
Idea, raised by Tim 22 Sepplatform and the library
Pricing per seat, with a credit limit and top-ups
Per seat, with a credit limit and top-ups; no number is set.
- The idea
- Tim's instinct of 22 Sep; the price needs its own research, and no number is set
Idea, raised by Tim 27 Sepon top: tests and proof
Mail checked in several email clients
Mail checked as it renders in real email clients, not only caught.
- Why
- A sign-up mail that arrives broken in one client is a lost user. Rendering checks show what each client really displays.
- The idea
- today mail is caught and its subject checked; rendering in Gmail and Outlook later, on approval
- What we know
- Today mail of 7 kinds is caught and its subject checked, on every runthe run at BookStack's development tip ·
- A free HTML check comes first; renders in Gmail and Outlook are paid and wait for approvalthe environments design ·
Idea, raised by Tim 27 Sepplatform and the library
Cheaper, steadier model calls across providers
Each model served by several providers: fall back when one is busy, pick the cheapest that is good enough.
- Why
- Model calls fail when one provider is busy. Spreading calls across providers keeps runs steady and cheaper, with a quality floor no call goes below.
- The idea
- each model is served by several providers; fall back when one is busy and pick the cheapest above a quality floor; researched, about 16% off a goal-checking run
- What we know
- Naming providers and setting a price ceiling came to about 16% off a goal-checking runprovider-routing research ·
- Batch mode runs at about half price, and its median batch finishes in 7 minutesprovider docs, read in research ·
Sources: the register on dev, 28 Sep 2026, the design notes of 21 to 28 Sep, the build log. Dates ahead are estimates; the board at /roadmap holds the same items by state.
An idea of your own? #
Tell us when you ask for a demo. We read each request and write back; the walkthrough is on BookStack.