Files
uhryniuk e06b371e2d Add autopilot-build skill: requirements-to-shipped-APK workflow
Codifies the ticket-writing, human-validation, wave-dispatch,
worktree-merge, emulator-verification, and APK-commit loop this repo's
UI-redesign and feedback batches were built with, written in ASD-STE100
so ticket files and subagent prompts stay unambiguous for a fresh
subagent with no conversation history.
2026-08-24 16:50:58 -05:00

16 KiB
Raw Permalink Blame History

name, description
name description
autopilot-build Turn a markdown requirements file into implemented, tested, emulator-verified, committed Flutter changes for this repo (rippr-flutter) — breaks requirements into independently-completable tickets, pauses for human approval, then dispatches implementation to isolated subagents, merges, verifies on the Android emulator, and commits a fresh APK.

Autopilot Build

Runs the same disciplined loop this repo's UI-redesign (docs/ui-redesign/) and post-launch-feedback (docs/feedback/) batches were built with, end to end, from a single requirements document. It is slow and thorough by design — every phase below exists because skipping it caused a real problem in a prior run of this loop.

Writing style: ASD-STE100 (Simplified Technical English)

Every ticket file, Outcome section, and subagent dispatch prompt this skill produces is written in ASD-STE100 — the controlled-language standard aerospace maintenance manuals use specifically to remove ambiguity for a reader with no shared context. That is exactly the constraint a fresh subagent with zero conversation history is under, so the standard's core rules apply directly, not as a style nicety:

  • One instruction, one sentence. Never join two instructions with "and" or a comma. "Remove the panel. Disconnect the cable." not "Remove the panel and disconnect the cable."
  • Active voice, imperative for instructions. "Add the constant to ride_map.dart." not "The constant should be added to ride_map.dart."
  • Keep sentences short (roughly 20 words or fewer). Split a long sentence at its natural joints rather than adding a subordinate clause.
  • Use one word for one meaning, every time. If a ticket calls it "the overflow menu" in one place, never call the same thing "the options menu" or "the kebab menu" later in the same ticket. Inconsistent naming is the single most common way a fresh-context reader misidentifies which code you mean.
  • Avoid long noun strings. "the HUD widget grid collision check" is four nouns stacked — prefer "the check that finds a collision in the HUD widget grid."
  • Avoid ambiguous pronouns. Repeat the noun instead of "it"/"this"/"that" whenever more than one candidate referent exists in the surrounding sentences — the exact failure mode a fresh subagent can't recover from is guessing which prior noun "it" meant.
  • Prefer concrete, specific verbs over vague ones. "Reject the move" not "handle the move"; "return the cached value" not "deal with the cached value."
  • State conditions before actions. "If the route has fewer than two waypoints, disable the toggle." not "Disable the toggle if the route has fewer than two waypoints" — STE100 puts the condition first so the reader isn't committed to an action before learning it might not apply.

This does not apply to direct quotations of existing code or user feedback (quote those verbatim, exactly as written) — only to the ticket prose you write around them.

Input

One argument: a path to a markdown requirements file (e.g. docs/FEEDBACK.md, or whatever the user hands you). If no path is given, ask for one — don't guess at requirements from conversation context alone; the file is the source of truth this whole run is grounded in.

Phase 0 — Sanity check before starting

  • flutter analyze clean and flutter test green. Record the passing count — every later phase's bar is "still clean, count only goes up from here."
  • git status clean (or confirm with the user what to do with any dirty state — don't silently work on top of someone else's uncommitted changes).
  • Confirm a remote exists if this run is expected to end in something pushable — this skill never pushes itself (see Phase 8), but say so up front if there's no remote.

Phase 1 — Break requirements into tickets

Read the requirements file in full. Read enough of the current codebase (grep, targeted file reads, and/or Explore-type subagents if the surface area is large or unfamiliar) to ground each requirement in real files, not guesses. For anything with real design ambiguity (a new data model, a rendering bug with no obvious cause, an architecture change), launch a Plan-type subagent to work through the approach before you commit to a ticket list — cheaper to redesign a paragraph now than a shipped ticket later.

Split the requirements into a small number of tickets, each one:

  • Independently completable — a fresh subagent with zero conversation history could read the ticket file alone and implement it correctly. This is the single most important property. If a ticket needs context only available in this conversation, the ticket is incomplete, not the subagent.
  • Scoped to avoid unnecessary file overlap with its siblings where possible. Where overlap is unavoidable (two tickets both need to touch the same screen), note the dependency explicitly — see Phase 2.

Do not write full ticket files yet. Recite the list to the user first: one line per ticket (title, one-sentence scope, rough size S/M/L). This mirrors the real worked example: five tickets were named and sized in one message, and only written out in full after the user confirmed the list. Skipping this step means discovering scope disagreements after hours of writing detail nobody asked for.

Phase 2 — Human validation (hard stop)

Do not proceed to Phase 3 without an explicit go-ahead on the ticket list. Use AskUserQuestion if there's a real fork to resolve (e.g. "should FB-04 and FB-05 be one ticket or two?"), otherwise just present the recited list and wait for the user's next message. Adjust scope, splitting, or sizing based on their response before moving on — this is the point where it's cheap to change direction.

Phase 3 — Deep-dive each ticket and write full ticket files

For tickets with non-trivial design questions still open, do another round of targeted investigation (fresh Explore/Plan subagent calls are fine here, run in parallel where the questions are independent of each other) until you have concrete, opinionated answers — not a menu of options — for every design decision the ticket will hand to an implementing subagent.

Write one ticket file per ticket into docs/<slug>/ (pick a short slug from the requirements doc's topic — docs/ui-redesign/, docs/feedback/ are the two existing precedents in this repo). Write every section in ASD-STE100 (see above). Each ticket file uses the same shape every ticket in this repo already uses:

# <ID> — <title>

**Depends on** <other ticket ids, or —> · **Size** S/M/L · **Status** Not started

## Goal
## Context       (exact current code excerpts, file paths, line numbers — quote, don't summarize)
## Design         (concrete, opinionated decisions — not options)
## Implementation (numbered steps)
## Acceptance criteria (checkboxes)
## Tests          (specific test names/cases to add, referencing existing test patterns in this repo)
## Risks
## Out of scope

Also write docs/<slug>/README.md: a table of ticket → size → depends-on → status, plus a dependency/dispatch-wave section (see Phase 4) and a one-paragraph working-method summary (implement → analyze/test clean → Outcome section → commit → status flip).

Commit the ticket docs as their own commit before dispatching anything.

Phase 4 — Sequence into dispatch waves

Two tickets that both edit the same file are a real merge-conflict risk when run in parallel isolated worktrees. Group tickets into waves:

  • Tickets with no file overlap with each other → same wave, dispatched in parallel.
  • A ticket that heavily edits a file another ticket already changed → its own later wave, dispatched only after the earlier one has merged (so its worktree branches from a base that already has the earlier change).
  • Light, disjoint-region overlap in the same file (e.g. two tickets touching different methods of the same screen) is an acceptable same-wave risk — it produced a small, manually-resolvable conflict in practice, not a blocker.

Write this wave plan into the ticket README before dispatching.

Phase 5 — Dispatch, implement, merge (per wave)

For each ticket in the current wave, launch an Agent call with isolation: "worktree" and the full ticket file contents inlined in the prompt (not just a path — a fresh agent should not have to go re-discover context this phase already gathered). Write the prompt itself in ASD-STE100 too — the dispatch prompt is the second place, after the ticket file, where ambiguity costs the most. The prompt must tell the subagent to:

  1. Read the inlined ticket, implement exactly what its Design/Implementation sections specify — it's already been reviewed, not something to redesign from scratch.
  2. Run flutter analyze (must stay clean) and flutter test (must stay green, count must not regress below the baseline from Phase 0/the previous wave).
  3. Append an ## Outcome section to its own ticket file: what shipped, any deviation from the design and why, the final test count. If an assumption in the ticket turns out wrong once real code is touched, say so honestly here rather than silently patching over it.
  4. Flip its own ticket's status line to Done — but not the shared README's table; that gets updated once after merging, by you, not by parallel subagents racing each other on the same file.
  5. Commit locally. Follow this repo's normal commit convention unless the user asked for something different for this run (e.g. no co-author trailer) — confirm which applies during Phase 2 if it's ambiguous, don't assume silently either way.
  6. State explicitly that Android-emulator verification is best-effort only at this stage (Phase 6 does the real device pass) — a subagent that gets stuck fighting emulator flakiness instead of finishing its actual ticket has failed the ticket.

Wait for each subagent in the wave to report back, then merge its branch:

git merge --no-ff worktree-agent-<id> -m "Merge <ticket-id>: <title>"

A conflict in the ticket .md file itself (both sides changed the status/Outcome header) is expected and mechanical — take the incoming side (git checkout --theirs docs/<slug>/<ticket>.md && git add ...), since it carries the Outcome section. A conflict in real code needs an actual read of both diffs — do not blindly pick one side; understand what each branch was trying to do and merge the intent, not just the text (see the worked example: two tickets each added a different, correct enhancement to the same function, and the right merge combined both rather than picking one).

After every merge: flutter analyze clean, flutter test green, count check. Only then move to the next ticket in the wave or the next wave. Update the ticket README's status table to Done for everything just merged, as its own small commit.

Phase 6 — Emulator verification pass (all waves merged)

This is not optional and not the same thing as the best-effort check individual subagents did — this is the real bar. Boot (or reuse, if already running and responsive) the project's emulator, launch the freshly-merged app, and walk every ticket's golden path for real, taking screenshots as evidence.

Known failure modes from prior runs, and how to handle them without losing an hour:

  • Coordinate scaling. Screenshots are shown to you scaled down (e.g. displayed at 900×2000 for a native 1080×2400 device) — the image caption states the exact ratio. Multiply both x and y by that ratio before issuing adb shell input tap. Getting this wrong looks like "the app isn't responding to taps" and wastes real time — it isn't a bug, it's a unit error. When in doubt, pixel-sample the raw screenshot file directly (PIL/Pillow in Python) to find a real on-screen coordinate rather than eyeballing the scaled-down preview.
  • ANR / "isn't responding" loops. Check host memory pressure first (vm_stat, sysctl vm.swapusage) before assuming an app bug — a memory-starved host produces exactly this symptom. Escalate in order: dismiss-and-wait once → force-stop and relaunch the app (adb shell am force-stop <pkg> then am start) → full emulator restart (adb emu kill + relaunch) → if the host itself is out of memory, say so and ask the user to free some rather than fighting it silently for another hour. com.rippr.port / com.rippr.MainActivity are this app's package/activity for direct am start (faster than a full flutter run rebuild when the APK is already installed and you just need a clean process).
  • Airplane mode / no network on a fresh emulator boot makes map tiles fail to load (renders as a checkerboard, not the app's real tiles) — check adb shell settings get global airplane_mode_on and adb shell cmd connectivity airplane-mode disable before concluding tiles are broken.
  • Stale app state (a leftover active recording, an old database) from a previous manual test session can make a screen look wrong for reasons unrelated to the current tickets — adb shell pm clear <pkg> before a verification pass if there's any doubt the state is clean.
  • Location-dependent features (anything reading GPS) need a granted permission and a mock fix on the emulator: adb shell pm grant <pkg> android.permission.ACCESS_FINE_LOCATION (and ACCESS_COARSE_LOCATION), then adb emu geo fix <lon> <lat> — note longitude first. A feature that reads location may take several seconds after granting/fixing before the first real fix arrives; don't conclude it's broken from a screenshot taken immediately after.

When the emulator pass finds a real bug (as opposed to an environment artifact above): fix it directly, add a regression test that fails without the fix and passes with it, re-run flutter analyze/flutter test, commit with a message that says what the emulator pass found and why the widget tests missed it, then re-verify on-device that the fix actually shows correctly. Do not mark a ticket's on-device verification done until you've seen the fixed behavior with your own eyes (a screenshot), not just inferred it from the test passing.

If the emulator is genuinely too unstable to get a real signal (host resource exhaustion that persists after the escalation steps above), say so plainly rather than claiming verification that didn't happen — this exact honesty call was made once already in this repo's history and is the correct call to make again if needed.

Phase 7 — Final commit: fresh APK

Once every ticket is merged, green, and emulator-verified:

flutter build apk --release
git add -f build/app/outputs/flutter-apk/app-release.apk
git commit -m "<summary of what this batch shipped>"

Force-add is required — build/ is gitignored, and this repo's own established precedent (see git log --oneline -- '*.apk') is to force-add the release APK directly into the repo rather than relying on CI artifacts or a release page, since there is no CI/release pipeline set up. Match that precedent unless the user has since set one up.

Phase 8 — Report back

Summarize: what shipped (ticket list with one line each), test count before → after, what the emulator pass confirmed (and any bugs it caught that the unit-test bar alone would have missed), and the final commit. State plainly whether anything was pushed (this skill never pushes on its own — say so, and say what's ready for the user to push if they want to).

What this skill is not for

A single small, obviously-scoped change (one file, one clear fix) doesn't need five phases of ceremony — just make the change, test it, done. This skill is for a requirements document with enough surface area that breaking it into independently verifiable pieces is itself valuable — the kind of input that would otherwise become one enormous, hard-to-review commit.