Files
samplez/tunie/research/RECOVERY_MODE.md
uhryniuk 4aa5da53d2 Commit tune maps and research so tunes are reachable from the phone
Removes the *.hex/maps_cache gitignore rule (explicit user call, reversing
the earlier no-redistribution stance) so the official TuneECU catalogue
maps, derived SAI/O2-delete composites, and the checksum/composition
tooling are actually available to pull up on a phone browser when using
the real TuneECU app. Also folds in tonight's KWP2000 fixes (TesterPresent
keep-alive, connect-failure cleanup, slow-init StartCommunication fix) and
the accumulated research docs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FP2GaxS9HkUdL5sLBnjKje
2026-08-27 01:34:05 -05:00

10 KiB

TuneECU's Recovery mode — traced (Aug 2026)

Started from menu_recovery in the decompile per docs/ROADMAP.md B2, to understand TuneECU's field-hardened recovery procedure (multi-year, multi-brand bug-fix history — Ducati, Aprilia Dorsoduro/Shiver, Walbro, 5DM/7SM ECUs all had recovery-specific fixes 2019-2021, per strings.xml's changelog) before this project ever attempts its own upload/write path (docs/ROADMAP.md B2/B4).

Priority is Triumph/Keihin recovery specifically, but the trace so far is mostly generic app architecture — the useful cross-brand lessons below are a side effect of that, not a separate investigation.

What "Recovery" actually is: not a separate feature, a relabeled one

MainActivity.java:30039: a single menu item (R.id.menu_prog) swaps its own label between "Reprogram" (menu_program) and "Recovery" (menu_recovery) based on a flag U9. Recovery is the same reprogram/write function as normal flashing — same code path (case R.id.menu_prog: at MainActivity.java:29129), gated by the same U9 condition that also swaps the label. There is no separate "Recovery" KWP2000 sequence sitting in its own function; whatever's different happens inside branches of the ordinary write path once U9 is true.

What decides U9 — and it's not live ECU status, it's file validation

Traced U9's assignment (MainActivity.java:9433, inside a large response/state-dispatch switch, case 10):

if (i8 == 0 && (k7 & 65280) == 2048) {   // (k7 & 0xFF00) == 0x0800
    z11 = true;
}
U9 = z11;

k7 comes from com.tuneecu.l.zc(bArrP8) (MainActivity.java:11011), where bArrP8 is the output of p8(str, true) — p8() is the same map file decoder this project already reverse-engineered (it's the function reconstruct_rom.py's docstring already cites for the unpack-directory format). So k7 isn't a live ECU status code at all — it's a property of the currently-loaded map file.

Opened l.zc() itself (l.java:6710) and it's doing work this project's own tooling already replicates:

  • Checks the same reversed magic-header pattern already documented in reference-maps/README.md (bytes[0:4] & 0xFF00FFE0 == 0x18001360) — returns an error code if it doesn't match.
  • Runs the same signature/Qd directory lookup table_map.py's resolve() already implements (sc() → s.a directory → Qd → c.a[Qd*48] calibration metadata).

So the practical meaning of U9 (show "Recovery" instead of "Reprogram") is: does the currently-open map file decode successfully and resolve to a valid calibration signature, of a type/size in a particular range (the 0x0800-high-byte check on whatever zc() returns) — not "the app detected the ECU is stuck via a live diagnostic read." This lines up exactly with the community-documented procedure found earlier (COMMUNITY_TUNING.md and forum research): "a matching map must definitely be opened, preferably a matching OEM map" before recovery works. The map has to be there and valid; the app isn't sensing the ECU's internal fault state, it's checking that you're prepared with a complete known-good image before it'll offer to write one.

The generalizable lesson (this is the actual cross-brand takeaway)

The design isn't "cleverly detect exactly where a failed transfer left off and resume it." It's "if you have a complete, valid, known-good image, do a full rewrite." That's a far more robust strategy than resume-logic — no need to reconstruct partial transfer state, no assumptions about where exactly a previous attempt died, just a full verified overwrite. This is the pattern worth carrying into any future tunie write-path work (docs/ROADMAP.md B4), regardless of ECU brand: always have a complete, validated base image ready before attempting a write, and treat "recovery" as "do the same full write again," not as a distinct clever-resume code path.

Counter-lesson, equally important: the changelog's multi-year, per-brand recovery bug list (5DM/7SM, Dorsoduro/Shiver 750, Walbro, Ducati — all separately broken and separately fixed, 2019-2021) shows that even with this simple, robust design, the implementation details differ enough per ECU family that a shared strategy still needed years of per-brand empirical patching to actually work reliably. "Full rewrite instead of resume" is the right architectural idea to borrow; assuming it works identically across ECU families without validating per-family is exactly the mistake that produced years of TuneECU's own bug list. Applies directly to this project: don't assume whatever works for Keihin generalizes to Sagem/Walbro/Bosch without separately checking each.

What happens after the confirmation dialog — traced through to the fork

Followed dialog id 43's positive-button click all the way through:

P9(...20, 3, 43) → confirm → s6.onClick() (MainActivity.java:6239) → H8(true, 43) → H8's case 27: case 43: (MainActivity.java:9527) shows a second confirmation dialog — the same standard write_caution warning a normal first-time reprogram shows, action code 12 → confirm again → H8's case 12: (MainActivity.java:9446), the actual fork:

case 12:
    if (!U9) {
        com.tuneecu.m.fg = true;      // normal path: flag read by the ordinary write routine
    } else {
        com.tuneecu.m.af();           // recovery path
    }
    break;

af() (m.java:6348) is short and concrete:

public static void af() {
    Zf = true; Yf = true; nf = true; Vf = true;   // mode flags for the write routine
    MainActivity.U9 = false;                       // clear the recovery flag
    Ue(d.MODE_NULL);                                // reset the connection state machine
    Jg.P9(null, null, Ig.ac(c.PLUG_SWITCH), 0, 20, 2, 15);  // prompt: cycle the ignition
    Qe();
    Ig.Yb(9600, true);          // reconnect at 9600 baud, with a control byte (3)
}

This independently confirms, from the code, exactly what the community procedure described from experience (COMMUNITY_TUNING.md/forum research): reset the connection, prompt the user to cycle the ignition switch (PLUG_SWITCH/UNPLUG_SWITCH are literal enum message keys for this), then reconnect — but at 9600 baud, not the normal K-line rate. That's a genuinely new, concrete, useful fact this project didn't have before: recovery mode reconnects at a different baud rate, strongly suggesting a slow/5-baud-style re-init rather than the normal fast init — which tunie already has support for (--init slow, ELM_INIT_SLOW in triumph.py), for the unrelated reason of "fast init timed out." The mechanism recovery leans on may be the same one tunie already implements for a different trigger condition.

Where the trace stops: what happens after the reconnect completes — the actual KWP2000 frames of the write/upload itself — isn't traced. The Zf/Yf/nf/Vf flags af() sets are presumably read by the same underlying write routine the normal fg-flag path uses, modifying its behavior (e.g. possibly skipping parts of SecurityAccess if a session is assumed already partially open) rather than being a wholly separate write implementation — consistent with the top-level finding that Recovery reuses the normal reprogram machinery rather than duplicating it. Tracing into that shared write routine itself is real further work, not attempted here.

Independent real-world confirmation (Aug 2026)

The file-based U9 detection this document traced — Recovery only becomes available when a valid map is loaded, not from a live ECU-fault check — is independently confirmed by real users, not just the static code trace. From the "TuneECU For Dummies" thread (triumphrat.net):

"you will NEED to have a map opened up in the TuneECU program when you go to reconnect to the bike or it will NOT initiate the recovery mode... try to connect, and then click OK when the recovery option is offered."

Matches the traced mechanism exactly — recovery is gated on a loaded, valid file, confirmed from both directions (static code and real usage). Also from the same source, a disconnect-ordering caution worth carrying into any future tunie write-path work: disconnect via the software menu before turning off ignition, not the other way around — one user reported turning off ignition first "closes the program mode on the ECU" incorrectly and the bike wouldn't start afterward until sorted out. Not independently traced in the code here, but a real reported failure mode worth respecting.

What's still untraced

  • The exact KWP2000 frames sent during the write itself — traced further, see WRITE_PATH.md: the shared routine both paths feed into (sc() → Fc() → a 5-baud slow-init bit-bang) is now documented there. That trace stops at the post-slow-init handoff (z.ec(...), unopened), which is the next link if this gets picked up again.
  • Whether there's also a live-ECU-side signal (e.g. a specific negative response during StartCommunication) that independently indicates a stuck programming session, separate from the file-based U9 check found here. Plausible — the community procedure's "cycle ignition, reconnect" step suggests the ECU's live response does matter somehow — but not confirmed from what's traced so far. Real next step if this gets picked up again: trace what happens between "ignition cycled, reconnect" and the recovery menu becoming available, since that's where a live-status check would live if one exists.
  • Byte-level meaning of the 0x0800 high-byte check on zc()'s return value — confirmed it's a classification of some kind (map type/size family), not confirmed exactly what distinguishes it from a normal map's classification.

Relevance to this project's own plans

  • Near-term (TUNING_IMPL_PLAN.md step 2): unaffected — still use TuneECU's app for the actual ROM dump, this doesn't change that.
  • Longer-term (docs/ROADMAP.md B2/B4, if tunie ever implements its own upload/write path): the "full rewrite over clever resume" pattern found here is directly actionable design guidance, and cheap to adopt — it's simpler to implement than resume logic would have been anyway. The per-brand-patching counter-lesson argues for validating any write path thoroughly against Keihin specifically before assuming it's solved in general, exactly matching this project's existing "get a spare ECU first" caution in safety.py.