timmy-talking-turd/ROADMAP.md

12 KiB
Raw Blame History

Timmy the Talking Turd — Product Roadmap

Canonical forge: https://forge.alexanderwhitestone.com/git/stackchain/timmy-talking-turd

North star

Make bowel logging fast enough to become a habit: take a photo, receive conservative observable-field suggestions, review or correct them, add symptoms yourself, and retain a private portable record. Timmy is playful during routine logging and calm/direct during safety escalation.

Product boundary

  • AI may suggest visible stool form, broad color, and image quality.
  • AI never diagnoses disease, identifies bleeding conclusively, infers pain/urgency/fever/vomiting, clears food, or suppresses deterministic red-flag escalation.
  • Model output remains provisional until the user confirms it.
  • Journal save and training contribution are separate consent decisions.
  • Hermes/Timmy remains release authority; human gates are limited to clinical/privacy review, beta consent, and RC approval.

Verified baseline at triage

  • Installable mobile-first PWA with manual Bristol logging, local ledger, JSON portability, privacy controls, history/calendar, summaries, and deterministic escalation.
  • Photo-first capture, explicit inference consent, strict model schema, confidence/abstention, and user confirmation.
  • Hosted and self-hosted OpenAI-compatible profiles with readiness reporting.
  • Real open-weight stool-photo spike completed: no content refusal, but the known Type 4 image was misclassified; production prefill remains blocked behind abstention and specialist-model evidence.
  • 22 unit/security tests, mobile green path, photo-first acceptance, syntax checks, and zero known npm vulnerabilities passed at triage.

Milestones

Milestone Target Outcome
M0 — Triage & Reproducible Baseline 2026-08-26 Canonical forge, CI, decisions, reproducible self-host bootstrap.
M1 — Sovereign Photo Intelligence 2026-09-16 Smooth camera-first UX, hardened sovereign inference, safety/privacy boundary.
M2 — Consented Dataset & Specialist Model 2026-10-15 Consent-governed corpus, evaluation harness, calibrated specialist model candidate.
M3 — Private Beta & Longitudinal Value 2026-11-15 User-owned longitudinal utility, grounded Timmy, bounded private beta.
M4 — Production Readiness & Release 2026-12-15 Security/load/ops evidence, RC manifest, rollback, explicit release approval.

Epic map

#1 — EPIC: Photo-first logging experience

#2 — EPIC: Sovereign vision inference

#3 — EPIC: Consented stool-image data flywheel

#4 — EPIC: Specialist Bristol classifier & calibration

#5 — EPIC: Clinical safety, privacy & security

#6 — EPIC: Journal intelligence & user-owned data

#7 — EPIC: Production operations, private beta & release

Critical sequence

  1. M0: preserve the green baseline, record the product boundary, and make the local worker reproducible.
  2. M1: harden capture and ingress, move inference to private GPU compute, prove real-stool schema/refusal behavior, and finish clinical/privacy language.
  3. M2: collect separately consented human-confirmed labels, run the 100-image reality check, then train only if the evidence says GO or NARROW.
  4. M3: prove repeated user value through explainable summaries, grounded Timmy chat, portable data, photo lifecycle controls, and a bounded beta.
  5. M4: attack the RC, exercise rollback, publish an evidence manifest, receive explicit approval, then refresh the film from the real release path.

Release gates

  • Contributor-held-out macro-F1 ≥ 0.80 across Bristol Types 17.
  • Precision ≥ 0.90 among non-abstained suggestions at the selected coverage.
  • Expected calibration error ≤ 0.05.
  • 100% valid Timmy schema or fail-closed response.
  • No inference-body or image-byte logging.
  • No training use without a versioned contribution receipt and tested withdrawal/deletion path.
  • Deterministic symptom escalation passes independently of model state.
  • Clean-checkout build, security/load suite, artifact hashes, rollback exercise, and explicit RC approval.

Definition of ready

  • One bounded outcome with measurable acceptance criteria.
  • Parent epic, milestone, labels, dependencies, privacy/safety boundary, and required evidence are present.
  • No real medical image is copied into an issue or Git history.

Definition of done

  • Focused tests and relevant full suite pass from a clean checkout.
  • PR links the issue and records reproducible evidence.
  • Privacy/safety impact is dispositioned.
  • Documentation and data/model provenance are updated.
  • Gitea issue is closed by the merged PR or explicitly deferred with rationale.

Planning policy

Gitea issues, milestones, and labels are the source of truth. Do not mirror this backlog into a second kanban unless workers actively consume that board. Keep contributor updates in the forge; surface only human gates or blocked decisions in chat.