feat: ship photo-first Timmy prototype and sovereign vision spike
7
.gitignore
vendored
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
node_modules/
|
||||
.playwright/
|
||||
npm-debug.log*
|
||||
__pycache__/
|
||||
*.pyc
|
||||
.env
|
||||
.env.*
|
||||
45
AI-EVIDENCE.md
Normal file
|
|
@ -0,0 +1,45 @@
|
|||
# AI Analysis Evidence — 2026-08-19
|
||||
|
||||
## Hosted path
|
||||
|
||||
- Transport: Timmy `/api/analyze` → local `hermes proxy` → Nous Portal OpenAI-compatible API.
|
||||
- Model: `stepfun/step-3.7-flash:free`.
|
||||
- Browser/app screenshot: `status: needs_user_input`, `isStool: false`.
|
||||
- Synthetic illustration: `status: needs_user_input`, `isStool: false`.
|
||||
|
||||
## Self-hosted open-weight path
|
||||
|
||||
- Runtime: locally compiled `llama.cpp`, commit `6d05498314db1b57f81c271080018aa2d0b89be9`.
|
||||
- API: loopback OpenAI-compatible `llama-server`.
|
||||
- Real input: metadata-stripped derivative of Wikimedia Commons **Human Feces Bristol Stool Chart Type 4**, SHA-256 `b3df757f75b83d20050675d14db9ac59d94e5065581748bc873bc5783782a63a`.
|
||||
- The source image is CC BY-SA 3.0 and explicitly labeled Type 4 by its source.
|
||||
|
||||
### SmolVLM-500M-Instruct Q8
|
||||
|
||||
- No content refusal.
|
||||
- Recognized stool, but did not produce a usable Bristol type.
|
||||
- Structured request: 25.606 seconds.
|
||||
- Resident memory: 908,420 KiB.
|
||||
- Verdict: admission/quality experiments only.
|
||||
|
||||
### SmolVLM2-2.2B-Instruct Q4_K_M
|
||||
|
||||
- No content refusal.
|
||||
- Valid schema.
|
||||
- `isStool: true`.
|
||||
- Color: `brown`.
|
||||
- Bristol prediction: Type 1; source label: Type 4.
|
||||
- Confidence: 0.5; quality: poor.
|
||||
- Timmy correctly returned `needs_user_input` instead of prefilling.
|
||||
- Cold structured request: 89.259 seconds.
|
||||
- Subsequent Timmy end-to-end request: 15.904 seconds.
|
||||
- Resident memory: 3,362,048 KiB.
|
||||
- Verdict: sovereign bootstrap worker; not an accurate production classifier.
|
||||
|
||||
## Positive UX path
|
||||
|
||||
The browser path—photo selection, compression, self-hosted readiness display, self-host-specific consent copy, API response, confidence display, Type/color prefill, confirmation, and preservation of zero-valued urgency/discomfort—passes with a deterministic positive fixture.
|
||||
|
||||
## Honest conclusion
|
||||
|
||||
The provider-moderation barrier is removed: a Timmy-controlled open-weight service accepts real stool imagery. The classification barrier remains: the generic VLM failed the known Type 4 label, so real prefilling must stay behind abstention while Timmy collects consented, human-confirmed labels for a specialist model. Full decision and sources: `research/SELF-HOSTED-STOOL-VISION.md`.
|
||||
59
PRODUCT.md
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
# Timmy the Talking Turd — Product Brief
|
||||
|
||||
## Promise
|
||||
|
||||
**Your intelligent pooping pal.** Timmy makes bowel tracking private, quick, funny, and useful without pretending a consumer photo is a medical diagnosis.
|
||||
|
||||
## Product wedge
|
||||
|
||||
Most stool trackers are sterile diaries. Most AI-health demos overclaim. Timmy combines:
|
||||
|
||||
1. a memorable character users actually return to;
|
||||
2. a ten-second Bristol-type log;
|
||||
3. an optional private photo attached to the entry;
|
||||
4. longitudinal, plain-language pattern summaries;
|
||||
5. fail-safe symptom escalation;
|
||||
6. user-owned export and one-tap deletion.
|
||||
|
||||
## MVP we can make real now
|
||||
|
||||
- Installable mobile-first PWA.
|
||||
- All records and optional photos remain in this browser's local storage.
|
||||
- Manual Bristol type 1–7 selection with clinically grounded groupings.
|
||||
- Color, urgency, discomfort, note, and urgent-symptom capture.
|
||||
- Calendar/history and deterministic pattern summary.
|
||||
- Photo-first capture that can send one compressed image, after explicit consent, to a configured multimodal AI provider.
|
||||
- AI prefills only the visible Bristol form and color, shows confidence, and requires confirmation.
|
||||
- Portable JSON export and delete-all control.
|
||||
- A Timmy voice that is playful about logging and serious about red flags.
|
||||
|
||||
## Honest AI boundary
|
||||
|
||||
The assisted-analysis release uses a general-purpose multimodal model to suggest only visible form and color. It is not clinically validated, never fills urgency, discomfort, symptoms, causes, treatment, or disease, and fails closed when confidence is low. Every suggestion requires user confirmation.
|
||||
|
||||
A production-grade image model still needs a consented, clinically labeled dataset; validation across lighting, toilets, skin tones, medications and diets; abstention when uncertain; privacy/security review; and likely medical-regulatory counsel if disease claims are ever introduced. Timmy should never clear a food, rule out disease, or replace professional care.
|
||||
|
||||
## Audience
|
||||
|
||||
- People managing constipation, diarrhea, IBS-like patterns, diet changes, medications, or a clinician-requested bowel diary.
|
||||
- Parents/caregivers need a distinct consent and child-safety flow; not in this MVP.
|
||||
|
||||
## Tone contract
|
||||
|
||||
- Funny: names, microcopy, streaks, and encouragement.
|
||||
- Never funny: blood, black/dark-red stool, severe/constant abdominal pain, vomiting, fever, inability to pass gas, unexplained weight loss, or emergencies.
|
||||
- Never say a food is "safe" from stool history.
|
||||
|
||||
## Evidence anchors
|
||||
|
||||
- Continence Health Australia: Bristol types 1–2 indicate constipation, 3–4 are healthy/easy to pass, and 5–7 indicate diarrhea or urgency: https://www.continence.org.au/about-incontinence/bowel-incontinence/bristol-stool-chart/
|
||||
- NHS: black/dark-red stool or bloody diarrhea warrants urgent help; non-stop or heavy bleeding is an emergency: https://www.nhs.uk/conditions/bleeding-from-the-bottom-rectal-bleeding/
|
||||
- NIDDK: constipation plus rectal bleeding, blood in stool, constant abdominal pain, inability to pass gas, vomiting, fever, or unintentional weight loss warrants prompt medical care: https://www.niddk.nih.gov/health-information/digestive-diseases/constipation/symptoms-causes
|
||||
|
||||
## Success metrics
|
||||
|
||||
- Median log completion under 20 seconds.
|
||||
- Seven-day retention and logs per active user.
|
||||
- Export completion and deletion reliability.
|
||||
- Zero diagnostic or food-clearance claims.
|
||||
- Red-flag flow always interrupts the playful UI with unambiguous action guidance.
|
||||
90
README.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
# Timmy the Talking Turd
|
||||
|
||||
A working mobile-first bowel diary with an optional **photo-first AI assist**: “Your intelligent pooping pal.”
|
||||
|
||||
## Run with the self-hosted open-weight path
|
||||
|
||||
```bash
|
||||
npm install
|
||||
|
||||
git clone --depth 1 https://github.com/ggml-org/llama.cpp /root/model-spikes/llama.cpp
|
||||
cmake -S /root/model-spikes/llama.cpp -B /root/model-spikes/llama.cpp/build -DCMAKE_BUILD_TYPE=Release
|
||||
cmake --build /root/model-spikes/llama.cpp/build -j4 --target llama-server
|
||||
|
||||
hf download ggml-org/SmolVLM2-2.2B-Instruct-GGUF \
|
||||
SmolVLM2-2.2B-Instruct-Q4_K_M.gguf \
|
||||
mmproj-SmolVLM2-2.2B-Instruct-Q8_0.gguf \
|
||||
--local-dir /root/model-spikes/models/smolvlm2-2.2b
|
||||
|
||||
TIMMY_MODEL_PORT=8080 scripts/run_selfhost_smolvlm.sh
|
||||
|
||||
# In another terminal
|
||||
TIMMY_VISION_PROFILE=selfhost npm start
|
||||
# open http://localhost:4173
|
||||
```
|
||||
|
||||
The selected bootstrap model is `SmolVLM2-2.2B-Instruct` using the official Apache-2.0 GGUF conversion. The default self-hosted endpoint is loopback-only at `http://127.0.0.1:8080/v1`. Override paths, host, port, threads, model ID, or endpoint with the `TIMMY_MODEL_*` and `TIMMY_VISION_*` environment variables.
|
||||
|
||||
## Run with a hosted provider
|
||||
|
||||
```bash
|
||||
hermes proxy start --provider nous --host 127.0.0.1 --port 8645
|
||||
npm start
|
||||
```
|
||||
|
||||
Hosted remains the default compatibility profile. To make it explicit, set `TIMMY_VISION_PROFILE=hosted`. Credentials stay server-side:
|
||||
|
||||
```bash
|
||||
TIMMY_VISION_BASE_URL=https://your-openai-compatible-provider/v1 \
|
||||
TIMMY_VISION_API_KEY=... \
|
||||
TIMMY_VISION_MODEL=your-vision-model \
|
||||
npm start
|
||||
```
|
||||
|
||||
Set `TIMMY_VISION_ENABLED=0` to disable uploads and retain manual-only operation.
|
||||
|
||||
## Verify
|
||||
|
||||
```bash
|
||||
npm test
|
||||
npm run test:ui # server required on port 4173
|
||||
npm run test:photo # mocked positive suggestion through the real browser UX
|
||||
npm audit --audit-level=high
|
||||
```
|
||||
|
||||
## What works
|
||||
|
||||
- Photo-first camera/file capture
|
||||
- Client-side image compression before transfer
|
||||
- Explicit consent before one-time provider analysis
|
||||
- Structured AI suggestions for **visible Bristol form and color only**
|
||||
- Confidence display, low-confidence abstention, and user confirmation
|
||||
- Urgency, discomfort, notes, and symptoms remain strictly user-reported
|
||||
- Manual logging that never uploads
|
||||
- Red-flag symptom escalation
|
||||
- Local browser ledger, calendar, pattern summary, JSON portability, and delete-all
|
||||
- Installable/offline PWA shell for the manual and saved-ledger paths
|
||||
|
||||
## What is deliberately not faked
|
||||
|
||||
The positive photo-prefill interface is covered by a deterministic provider fixture. A live self-hosted SmolVLM2-2.2B model was exercised against a CC-licensed real Type 4 stool photograph. It accepted the content and identified stool, but predicted the wrong type with low confidence; Timmy correctly abstained. **The open model removes provider moderation from the path, but its Bristol accuracy is not clinically established.**
|
||||
|
||||
Timmy does not diagnose disease, identify bleeding, infer pain/urgency, recommend treatment, or clear foods. The selected VLM is a bootstrap worker, not the final classifier. See [research/SELF-HOSTED-STOOL-VISION.md](research/SELF-HOSTED-STOOL-VISION.md) for the measured receipts, hardware gate, consented-data pipeline, specialist-classifier plan, and sources.
|
||||
|
||||
## Architecture
|
||||
|
||||
```text
|
||||
Browser PWA
|
||||
├── app.js photo-first UX, local persistence, compression
|
||||
├── src/domain.js tested health/safety and summary rules
|
||||
├── src/analysis.js strict AI schema, validation, visual-only merge
|
||||
├── server.mjs static server + bounded /api/analyze route
|
||||
├── src/vision-service.js server-side OpenAI-compatible provider adapter
|
||||
├── src/vision-config.js hosted/self-hosted profiles and readiness probe
|
||||
├── scripts/run_selfhost_smolvlm.sh
|
||||
├── scripts/ingest_training_photo.py
|
||||
├── localStorage saved ledger and attached photos
|
||||
└── llama.cpp / hosted API selected per server-side profile
|
||||
```
|
||||
|
||||
The server accepts JPEG/PNG/WebP data URLs up to 4 MB, uses a bounded 60-second hosted or 120-second self-hosted timeout, never accepts an API key from browser input, does not write inference images to disk, and returns only validated visual suggestions. The separate dataset ingester runs only for explicit, consented contributions.
|
||||
127
app.js
Normal file
|
|
@ -0,0 +1,127 @@
|
|||
import { bucketForBristolType, buildTimmySummary, detectUrgentFlags, exportLedger, importLedger, photoQualityMessage, sanitizeEntry } from './src/domain.js';
|
||||
import { mergeVisualSuggestion } from './src/analysis.js';
|
||||
|
||||
const STORE = 'timmy-ledger-v1';
|
||||
const app = document.querySelector('#app');
|
||||
let entries = loadEntries();
|
||||
let view = 'home';
|
||||
let photoDataUrl = '';
|
||||
let photoHint = '';
|
||||
let aiSuggestion = null;
|
||||
let visionStatus = null;
|
||||
let photoFirstMode = 'pick';
|
||||
const draft = () => ({ bristolType: 4, color: 'brown', urgency: 0, discomfort: 0, note: '', symptoms: {} });
|
||||
let form = draft();
|
||||
|
||||
function esc(value='') { return String(value).replace(/[&<>'"]/g, c => ({'&':'&','<':'<','>':'>',"'":''','"':'"'}[c])); }
|
||||
function loadEntries() { try { return JSON.parse(localStorage.getItem(STORE) || '[]'); } catch { return []; } }
|
||||
function saveEntries() { localStorage.setItem(STORE, JSON.stringify(entries)); }
|
||||
function formatDate(value) { return new Intl.DateTimeFormat(undefined,{month:'short',day:'numeric',hour:'numeric',minute:'2-digit'}).format(new Date(value)); }
|
||||
function toast(message) { const node=document.createElement('div');node.className='toast';node.textContent=message;document.body.append(node);setTimeout(()=>node.remove(),2400); }
|
||||
|
||||
function shell(content) {
|
||||
app.innerHTML = `<header class="topbar"><div class="brand"><img src="/assets/timmy.svg" alt="Timmy mascot"><div class="brand-copy"><strong>Timmy</strong><span>the Talking Turd</span></div></div><div class="privacy-chip">🔒 Local ledger</div></header>${content}${nav()}`;
|
||||
bindGlobal();
|
||||
}
|
||||
function nav(){return `<nav class="bottom-nav" aria-label="Primary"><button class="nav-btn ${view==='home'?'active':''}" data-view="home"><b>⌂</b>Home</button><button class="nav-btn ${view==='calendar'?'active':''}" data-view="calendar"><b>▦</b>Calendar</button><button class="nav-btn ${view==='timmy'?'active':''}" data-view="timmy"><b>◉</b>Ask Timmy</button><button class="nav-btn ${view==='privacy'?'active':''}" data-view="privacy"><b>⌁</b>Privacy</button></nav>`}
|
||||
function bindGlobal(){
|
||||
document.querySelectorAll('[data-view]').forEach(btn=>btn.onclick=()=>{view=btn.dataset.view;render()});
|
||||
document.querySelectorAll('[data-log]').forEach(btn=>btn.onclick=openLogger);
|
||||
document.querySelectorAll('[data-scan]').forEach(btn=>btn.onclick=openPhotoFirst);
|
||||
}
|
||||
|
||||
function recentList(limit=5){
|
||||
if(!entries.length)return `<div class="empty"><img src="/assets/timmy.svg" alt=""><h3>Quiet bowl, clean slate.</h3><p>Your first log takes about ten seconds.</p></div>`;
|
||||
return entries.slice().sort((a,b)=>new Date(b.occurredAt)-new Date(a.occurredAt)).slice(0,limit).map(e=>`<article class="entry"><div class="type-dot">T${e.bristolType}</div><div><strong>${formatDate(e.occurredAt)}</strong><span>${esc(e.color)} · urgency ${e.urgency}/4 · discomfort ${e.discomfort}/4${e.note?` · ${esc(e.note)}`:''}</span></div><span class="bucket bucket-${bucketForBristolType(e.bristolType)}">${bucketForBristolType(e.bristolType)}</span></article>`).join('');
|
||||
}
|
||||
function thisWeek(){const now=Date.now(),week=7*864e5;return entries.filter(e=>now-new Date(e.occurredAt).getTime()<week).length}
|
||||
function currentStreak(){const dates=new Set(entries.map(e=>e.occurredAt.slice(0,10)));let n=0,d=new Date();while(dates.has(d.toISOString().slice(0,10))){n++;d.setDate(d.getDate()-1)}return n}
|
||||
|
||||
function home(){
|
||||
shell(`<main><section class="hero"><span class="eyebrow">Your intelligent pooping pal</span><h1>Snap first.<br>Timmy fills the form.</h1><p class="lead">Take a private photo. AI suggests the visible Bristol form and color; you confirm it, then add the things a camera cannot know.</p><div class="hero-actions"><button class="btn btn-primary btn-scan" data-scan>📷 Analyze a photo</button><button class="btn btn-secondary" data-log>Log manually</button></div><p class="hero-foot">AI photo mode is optional. Your saved ledger stays in this browser.</p></section><section class="section"><div class="stats"><div class="stat"><strong>${thisWeek()}</strong><span>THIS WEEK</span></div><div class="stat"><strong>${currentStreak()}</strong><span>DAY STREAK</span></div><div class="stat"><strong>${entries.length}</strong><span>ALL LOGS</span></div></div></section><section class="section"><div class="section-head"><div><span class="eyebrow">Timmy noticed</span><h2>Your pattern</h2></div></div><div class="card summary-card"><img src="/assets/timmy.svg" alt=""><p>${esc(buildTimmySummary(entries))}</p></div></section><section class="section"><div class="section-head"><div><span class="eyebrow">Recent business</span><h2>Your logs</h2></div>${entries.length?'<button class="btn btn-ghost" data-view="calendar">See all</button>':''}</div><div class="card">${recentList(4)}</div></section></main>`);
|
||||
}
|
||||
|
||||
function calendar(){
|
||||
const now=new Date(),year=now.getFullYear(),month=now.getMonth(),first=new Date(year,month,1),days=new Date(year,month+1,0).getDate();
|
||||
const counts={};entries.forEach(e=>{const d=new Date(e.occurredAt);if(d.getFullYear()===year&&d.getMonth()===month)counts[d.getDate()]=(counts[d.getDate()]||0)+1});
|
||||
const blanks=Array(first.getDay()).fill('<div class="day blank"></div>').join('');
|
||||
const boxes=Array.from({length:days},(_,i)=>`<button class="day ${counts[i+1]?'has-log':''}" title="${counts[i+1]||0} logs">${i+1}</button>`).join('');
|
||||
shell(`<main><div class="page-title"><span class="eyebrow">The poop calendar</span><h1>${now.toLocaleString(undefined,{month:'long'})}</h1><p>A calm view of frequency and form. One unusual day is not a verdict.</p></div><section class="card"><div class="calendar">${['S','M','T','W','T','F','S'].map(x=>`<div class="cal-head">${x}</div>`).join('')}${blanks}${boxes}</div></section><section class="section"><div class="section-head"><h2>All entries</h2><button class="btn btn-primary" data-log>+ Add</button></div><div class="card">${recentList(100)}</div></section></main>`);
|
||||
}
|
||||
|
||||
function timmy(){
|
||||
const reply=buildTimmySummary(entries);
|
||||
shell(`<main><div class="page-title"><span class="eyebrow">Pattern pal, not a doctor</span><h1>Ask Timmy</h1><p>Timmy answers from the records on this device. He never diagnoses or clears a food.</p></div><section class="card"><div class="chat" id="chat"><div class="bubble timmy">Hey, bowel buddy. I can summarize your recent form and frequency or explain what this prototype stores.</div><div class="bubble timmy">${esc(reply)}</div></div><div class="prompt-row section"><button class="prompt" data-prompt="pattern">What’s my pattern?</button><button class="prompt" data-prompt="privacy">Where are my photos?</button><button class="prompt" data-prompt="food">Can I eat Taco Bell?</button></div></section><section class="section card"><h3>Timmy’s hard boundary</h3><p class="fine">If you report blood, black or dark-red stool, severe or constant abdominal pain, vomiting, fever, or inability to pass gas, Timmy stops joking and tells you to seek medical care.</p></section></main>`);
|
||||
document.querySelectorAll('[data-prompt]').forEach(btn=>btn.onclick=()=>chatReply(btn));
|
||||
}
|
||||
function chatReply(btn){
|
||||
const chat=document.querySelector('#chat'),kind=btn.dataset.prompt;
|
||||
const q={pattern:'What’s my pattern?',privacy:'Where are my photos?',food:'Can I eat Taco Bell?'}[kind];
|
||||
const a={pattern:buildTimmySummary(entries),privacy:'Your saved ledger and photos stay in this browser. If you explicitly use Analyze a photo, one compressed copy is sent to the configured AI provider for that analysis and is not stored by Timmy’s server.',food:'That call is yours. A stool diary cannot clear a restaurant or prove a food is safe. Log what happens and look for repeated patterns.'}[kind];
|
||||
chat.insertAdjacentHTML('beforeend',`<div class="bubble user">${q}</div><div class="bubble timmy">${esc(a)}</div>`);
|
||||
}
|
||||
|
||||
function privacy(){
|
||||
shell(`<main><div class="page-title"><span class="eyebrow">Private by design</span><h1>Your poop. Your phone.</h1><p>This prototype has no account, analytics, ad tracker, or server database.</p></div><section class="card privacy-list"><div class="privacy-item"><b>⌂</b><div><h3>Stored locally</h3><p>Saved entries and optional photos live in this browser’s local storage.</p></div></div><div class="privacy-item"><b>⇩</b><div><h3>Portable</h3><p>Export a readable JSON file. Import it in another copy of Timmy.</p></div></div><div class="privacy-item"><b>◎</b><div><h3>AI only when you ask</h3><p>Manual logging never uploads. Analyze a photo sends one compressed copy to the configured AI provider after consent; Timmy’s server does not save it.</p></div></div></section><section class="section card"><h2>Data controls</h2><div class="row"><button class="btn btn-primary" id="export">Export JSON</button><label class="btn btn-secondary" for="import">Import JSON</label><input class="hidden" type="file" id="import" accept="application/json"></div><p class="fine section">Exports can contain sensitive health notes and photos. Store them somewhere you trust.</p></section><section class="section card source-list"><h2>Health sources</h2><p class="fine">The Bristol groupings and urgent-symptom copy are grounded in public clinical guidance.</p><p><a target="_blank" rel="noreferrer" href="https://www.continence.org.au/about-incontinence/bowel-incontinence/bristol-stool-chart/">Continence Health Australia</a></p><p><a target="_blank" rel="noreferrer" href="https://www.nhs.uk/conditions/bleeding-from-the-bottom-rectal-bleeding/">NHS rectal bleeding guidance</a></p><p><a target="_blank" rel="noreferrer" href="https://www.niddk.nih.gov/health-information/digestive-diseases/constipation/symptoms-causes">NIDDK constipation guidance</a></p></section><section class="section card danger-zone"><h2>Delete everything</h2><p class="fine">Permanently removes Timmy’s local ledger from this browser.</p><button class="btn btn-danger" id="delete-all">Delete all local data</button></section></main>`);
|
||||
document.querySelector('#export').onclick=exportData;document.querySelector('#import').onchange=importData;document.querySelector('#delete-all').onclick=deleteData;
|
||||
}
|
||||
function exportData(){const blob=new Blob([exportLedger(entries)],{type:'application/json'}),a=document.createElement('a');a.href=URL.createObjectURL(blob);a.download='timmy-ledger.json';a.click();URL.revokeObjectURL(a.href);toast('Export created');}
|
||||
async function importData(e){try{const text=await e.target.files[0].text();entries=importLedger(text);saveEntries();render();toast('Ledger imported')}catch(err){toast(err.message)}}
|
||||
function deleteData(){if(confirm('Delete every local Timmy entry and photo? This cannot be undone.')){entries=[];localStorage.removeItem(STORE);render();toast('Local ledger deleted')}}
|
||||
|
||||
function openPhotoFirst(){form=draft();photoDataUrl='';photoHint='';aiSuggestion=null;visionStatus=null;showPhotoFirst('pick');loadVisionStatus()}
|
||||
async function loadVisionStatus(){
|
||||
try{const response=await fetch('/api/vision-status',{headers:{accept:'application/json'}});visionStatus=response.ok?await response.json():{enabled:false,providerReady:false}}
|
||||
catch{visionStatus={enabled:false,providerReady:false}}
|
||||
if(document.querySelector('.scan-sheet')&&['pick','ready'].includes(photoFirstMode))showPhotoFirst(photoFirstMode)
|
||||
}
|
||||
function visionStatusHtml(){
|
||||
if(!visionStatus)return '<div class="model-status checking">◌ Checking the vision worker…</div>';
|
||||
if(visionStatus.profile==='selfhost'&&visionStatus.providerReady)return `<div class="model-status ready">● Self-hosted model ready · ${esc(visionStatus.model)}</div>`;
|
||||
if(visionStatus.providerReady)return `<div class="model-status ready">● Vision provider ready · ${esc(visionStatus.model)}</div>`;
|
||||
return '<div class="model-status offline">○ Vision worker offline · manual logging is still available</div>';
|
||||
}
|
||||
function photoFirstBody(mode,error=''){
|
||||
if(mode==='pick')return `${visionStatusHtml()}<div class="scan-hero"><img src="/assets/timmy.svg" alt="Timmy"><h3>One photo. Two useful suggestions.</h3><p>Timmy can suggest the visible Bristol form and color. A camera cannot know urgency, pain, symptoms, or a diagnosis.</p></div><label class="photo-capture" for="ai-photo"><b>📷</b><strong>Take or choose a photo</strong><span>JPEG, PNG, or WebP · compressed before analysis</span><input id="ai-photo" type="file" accept="image/*" capture="environment"></label><button class="btn btn-ghost btn-wide section" id="manual-from-scan">Continue without AI</button>`;
|
||||
if(mode==='ready'){const processingCopy=visionStatus?.profile==='selfhost'?'Timmy’s server does not save it. The compressed copy stays on Timmy’s self-hosted model server.':'Timmy’s server does not save it. Your configured AI provider processes it under that provider’s terms.';return `${visionStatusHtml()}<img class="photo-preview scan-preview" src="${photoDataUrl}" alt="Photo awaiting AI analysis"><p class="quality-note">${esc(photoHint)}</p><div class="consent-card"><label class="check"><input id="ai-consent" type="checkbox"><span><strong>Send this compressed copy for one-time AI analysis.</strong><br>${processingCopy}</span></label></div><button class="btn btn-primary btn-wide" id="analyze-photo" disabled>Analyze visible form + color</button><button class="btn btn-ghost btn-wide section" id="retake-photo">Use another photo</button>`;}
|
||||
if(mode==='analyzing')return `<div class="analyzing"><img src="/assets/timmy.svg" alt="Timmy"><div class="spinner" aria-hidden="true"></div><h3>Timmy is looking at form and color…</h3><p>Not symptoms. Not disease. Not whether Taco Bell was a strategic error.</p></div>`;
|
||||
if(mode==='error')return `<div class="scan-result needs-input"><b>↻</b><h3>Timmy couldn’t analyze that safely.</h3><p>${esc(error||'Continue manually or try a clearer photo.')}</p></div><button class="btn btn-primary btn-wide" id="manual-from-scan">Fill it out manually</button><button class="btn btn-ghost btn-wide section" id="retake-photo">Try another photo</button>`;
|
||||
if(aiSuggestion?.status==='suggestion')return `<div class="scan-result"><span class="ai-badge">AI SUGGESTION · ${Math.round(aiSuggestion.confidence*100)}% CONFIDENCE</span><div class="suggestion-pair"><div><small>BRISTOL FORM</small><strong>Type ${aiSuggestion.bristolType}</strong></div><div><small>VISIBLE COLOR</small><strong>${esc(aiSuggestion.color)}</strong></div></div><p>${esc(aiSuggestion.observations||'Visual match found.')}</p><p class="fine">${esc(aiSuggestion.warning)}</p></div><button class="btn btn-primary btn-wide" id="use-suggestion">Use these suggestions</button><button class="btn btn-ghost btn-wide section" id="manual-from-scan">Review everything manually</button>`;
|
||||
return `<div class="scan-result needs-input"><b>?</b><h3>No confident match.</h3><p>${esc(aiSuggestion?.reason||'The image was too uncertain to prefill safely.')}</p></div><button class="btn btn-primary btn-wide" id="manual-from-scan">Choose the form yourself</button><button class="btn btn-ghost btn-wide section" id="retake-photo">Try another photo</button>`;
|
||||
}
|
||||
function showPhotoFirst(mode='pick',error=''){
|
||||
photoFirstMode=mode;
|
||||
document.querySelector('.sheet-backdrop')?.remove();const wrap=document.createElement('div');wrap.className='sheet-backdrop';wrap.innerHTML=`<section class="sheet scan-sheet" role="dialog" aria-modal="true" aria-labelledby="scan-title"><div class="sheet-handle"></div><div class="sheet-header"><div><span class="eyebrow">Photo-first log</span><h2 id="scan-title">${mode==='result'?'Review Timmy’s suggestion':mode==='analyzing'?'Analyzing privately':'Start with the camera'}</h2></div><button class="icon-btn" id="close-sheet" aria-label="Close">×</button></div>${photoFirstBody(mode,error)}</section>`;document.body.append(wrap);document.querySelector('#close-sheet').onclick=()=>wrap.remove();wrap.onclick=e=>{if(e.target===wrap)wrap.remove()};
|
||||
const file=document.querySelector('#ai-photo');if(file)file.onchange=handleAiPhoto;
|
||||
const consent=document.querySelector('#ai-consent'),analyze=document.querySelector('#analyze-photo');if(consent&&analyze)consent.onchange=()=>analyze.disabled=!consent.checked||visionStatus?.providerReady===false;if(analyze)analyze.onclick=runAiAnalysis;
|
||||
document.querySelector('#retake-photo')?.addEventListener('click',()=>{photoDataUrl='';photoHint='';aiSuggestion=null;showPhotoFirst('pick')});
|
||||
document.querySelector('#manual-from-scan')?.addEventListener('click',()=>{aiSuggestion=null;showLogStep(1)});
|
||||
document.querySelector('#use-suggestion')?.addEventListener('click',()=>{form=mergeVisualSuggestion(form,aiSuggestion);showLogStep(1)});
|
||||
}
|
||||
async function handleAiPhoto(e){const file=e.target.files[0];if(!file)return;try{const result=await compressPhoto(file);photoDataUrl=result.dataUrl;photoHint=photoQualityMessage(result);showPhotoFirst('ready')}catch{showPhotoFirst('error','That image could not be read. Try another photo.')}}
|
||||
async function runAiAnalysis(){showPhotoFirst('analyzing');try{const response=await fetch('/api/analyze',{method:'POST',headers:{'content-type':'application/json'},body:JSON.stringify({imageDataUrl:photoDataUrl,consent:true})});const data=await response.json();if(!response.ok)throw new Error(data.error||'AI analysis is unavailable.');aiSuggestion=data;showPhotoFirst('result')}catch(error){showPhotoFirst('error',error.message||'AI analysis is unavailable. Continue manually.')}}
|
||||
|
||||
function openLogger(){form=draft();photoDataUrl='';photoHint='';aiSuggestion=null;showLogStep(1)}
|
||||
function showLogStep(step){
|
||||
document.querySelector('.sheet-backdrop')?.remove();
|
||||
const wrap=document.createElement('div');wrap.className='sheet-backdrop';wrap.innerHTML=`<section class="sheet" role="dialog" aria-modal="true" aria-labelledby="log-title"><div class="sheet-handle"></div><div class="sheet-header"><div><span class="eyebrow">Step ${step} of 3</span><h2 id="log-title">${step===1?'Pick the closest form':step===2?'Add useful context':'Safety check'}</h2></div><button class="icon-btn" id="close-sheet" aria-label="Close">×</button></div><div class="progress"><i style="width:${step*33.34}%"></i></div>${stepBody(step)}</section>`;document.body.append(wrap);
|
||||
document.querySelector('#close-sheet').onclick=()=>wrap.remove();wrap.onclick=e=>{if(e.target===wrap)wrap.remove()};bindStep(step);
|
||||
}
|
||||
function stepBody(step){
|
||||
if(step===1)return `${aiSuggestion?.status==='suggestion'?`<div class="ai-prefill"><span class="ai-badge">AI PREFILLED</span><strong>Type ${aiSuggestion.bristolType} · ${esc(aiSuggestion.color)}</strong><small>You’re in charge—tap any type to correct it.</small></div>`:'<p class="fine">Choose the closest match yourself. Timmy never treats a suggestion as fact.</p>'}<div class="choice-grid">${[[1,'Hard separate lumps'],[2,'Lumpy sausage'],[3,'Cracked sausage'],[4,'Smooth and soft'],[5,'Soft blobs'],[6,'Mushy pieces'],[7,'Entirely liquid']].map(([n,d])=>`<button class="bristol ${form.bristolType===n?'selected':''}" data-type="${n}"><strong>Type ${n}</strong><span>${d}</span></button>`).join('')}</div><button class="btn btn-primary btn-wide section" id="next">Confirm + add details →</button>`;
|
||||
if(step===2)return `<label class="field"><span class="field-label">Color</span><select class="input" id="color"><option>brown</option><option>green</option><option>yellow</option><option>pale</option><option>red</option><option>black</option></select></label><label class="field"><span class="field-label">Urgency</span><div class="range-row"><input id="urgency" type="range" min="0" max="4" value="${form.urgency}"><output class="range-val">${form.urgency}</output></div></label><label class="field"><span class="field-label">Discomfort</span><div class="range-row"><input id="discomfort" type="range" min="0" max="4" value="${form.discomfort}"><output class="range-val">${form.discomfort}</output></div></label><label class="field"><span class="field-label">Note (optional)</span><textarea class="input" id="note" maxlength="500" placeholder="Meal, medicine, travel, stress…">${esc(form.note)}</textarea></label><label class="photo-drop btn" for="photo">📷 Add a private photo (optional)<input id="photo" type="file" accept="image/*" capture="environment"></label>${photoDataUrl?`<img class="photo-preview" src="${photoDataUrl}" alt="Private entry preview">`:''}${photoHint?`<p class="fine">${esc(photoHint)}</p>`:''}${aiSuggestion?.status==='suggestion'?'<p class="fine">AI suggested only form and visible color. Urgency, discomfort, notes, and symptoms must come from you.</p>':'<p class="fine">Manual-mode photos stay in this browser and receive quality checks only.</p>'}<div class="row section"><button class="btn btn-ghost" id="back">← Back</button><button class="btn btn-primary" id="next">Safety check →</button></div>`;
|
||||
const urgent=detectUrgentFlags(form.symptoms);return `<p class="fine">Select anything you have now. This is where Timmy stops joking.</p><div class="symptoms">${[['blood','Blood in stool or rectal bleeding'],['blackOrDarkRed','Black or dark-red stool'],['severePain','Severe or constant abdominal pain'],['vomiting','Vomiting'],['fever','Fever'],['cannotPassGas','Unable to pass gas']].map(([k,l])=>`<label class="check"><input type="checkbox" data-symptom="${k}" ${form.symptoms[k]?'checked':''}><span>${l}</span></label>`).join('')}</div><div id="urgent-box">${urgent.urgent?alertHtml(urgent.message):''}</div><div class="row section"><button class="btn btn-ghost" id="back">← Back</button><button class="btn btn-primary" id="save">Save private log</button></div><p class="fine">Not medical advice. Heavy or nonstop bleeding, fainting, or severe worsening symptoms can be an emergency—call local emergency services.</p>`;
|
||||
}
|
||||
function alertHtml(msg){return `<div class="alert"><strong>Pause and get medical help.</strong><p>${esc(msg)}</p></div>`}
|
||||
function bindStep(step){
|
||||
if(step===1){document.querySelectorAll('[data-type]').forEach(b=>b.onclick=()=>{form.bristolType=Number(b.dataset.type);aiSuggestion=null;showLogStep(1)});document.querySelector('#next').onclick=()=>showLogStep(2)}
|
||||
if(step===2){const color=document.querySelector('#color');color.value=form.color;color.onchange=()=>form.color=color.value;['urgency','discomfort'].forEach(k=>{const n=document.querySelector('#'+k);n.oninput=()=>{form[k]=Number(n.value);n.nextElementSibling.value=n.value}});document.querySelector('#note').oninput=e=>form.note=e.target.value;document.querySelector('#photo').onchange=handlePhoto;document.querySelector('#back').onclick=()=>showLogStep(1);document.querySelector('#next').onclick=()=>showLogStep(3)}
|
||||
if(step===3){document.querySelectorAll('[data-symptom]').forEach(c=>c.onchange=()=>{form.symptoms[c.dataset.symptom]=c.checked;document.querySelector('#urgent-box').innerHTML=detectUrgentFlags(form.symptoms).urgent?alertHtml(detectUrgentFlags(form.symptoms).message):''});document.querySelector('#back').onclick=()=>showLogStep(2);document.querySelector('#save').onclick=saveLog}
|
||||
}
|
||||
async function handlePhoto(e){const file=e.target.files[0];if(!file)return;try{const result=await compressPhoto(file);photoDataUrl=result.dataUrl;photoHint=photoQualityMessage(result);showLogStep(2)}catch{photoHint='That image could not be read. Try another photo.';showLogStep(2)}}
|
||||
function compressPhoto(file){return new Promise((resolve,reject)=>{const img=new Image(),url=URL.createObjectURL(file);img.onload=()=>{const scale=Math.min(1,1200/Math.max(img.width,img.height)),canvas=document.createElement('canvas');canvas.width=Math.round(img.width*scale);canvas.height=Math.round(img.height*scale);const ctx=canvas.getContext('2d');ctx.drawImage(img,0,0,canvas.width,canvas.height);const sample=ctx.getImageData(0,0,Math.min(canvas.width,120),Math.min(canvas.height,120)).data;let total=0;for(let i=0;i<sample.length;i+=4)total+=(sample[i]+sample[i+1]+sample[i+2])/3;URL.revokeObjectURL(url);resolve({dataUrl:canvas.toDataURL('image/jpeg',.7),width:img.width,height:img.height,brightness:total/(sample.length/4)/255})};img.onerror=()=>{URL.revokeObjectURL(url);reject()};img.src=url})}
|
||||
function saveLog(){const result=detectUrgentFlags(form.symptoms);const entry=sanitizeEntry({...form,photoDataUrl});try{entries.push(entry);saveEntries()}catch{entry.photoDataUrl='';entries[entries.length-1]=entry;saveEntries();toast('Log saved, but the photo was too large for browser storage')}document.querySelector('.sheet-backdrop')?.remove();view='home';render();toast(result.urgent?'Saved. Please follow the medical-care alert.':'Private log saved')}
|
||||
function render(){({home,calendar,timmy,privacy}[view]||home)()}
|
||||
|
||||
render();
|
||||
if('serviceWorker' in navigator)navigator.serviceWorker.register('/service-worker.js').catch(()=>{});
|
||||
BIN
artifacts/home-mobile.png
Normal file
|
After Width: | Height: | Size: 389 KiB |
BIN
artifacts/photo-first-prefill-mobile.png
Normal file
|
After Width: | Height: | Size: 173 KiB |
BIN
artifacts/photo-first-result-mobile.png
Normal file
|
After Width: | Height: | Size: 386 KiB |
BIN
artifacts/red-flag-mobile.png
Normal file
|
After Width: | Height: | Size: 195 KiB |
BIN
artifacts/selfhost-photo-first-mobile.png
Normal file
|
After Width: | Height: | Size: 318 KiB |
BIN
artifacts/timmy-chat-mobile.png
Normal file
|
After Width: | Height: | Size: 221 KiB |
1
assets/icon-192.svg
Normal file
|
|
@ -0,0 +1 @@
|
|||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 192 192"><rect width="192" height="192" rx="46" fill="#f5c95b"/><path d="M43 144c-5-20 7-36 26-40-10-10-7-24 7-31-3-13 7-26 21-26 18 0 27 16 21 30 17 5 24 22 15 34 17 7 24 22 17 36-10 22-94 22-107-3Z" fill="#75452f" stroke="#3f251e" stroke-width="6"/><path d="M66 52c18-20 61-22 76 6-20-7-46-6-70 6Z" fill="#28a9a1" stroke="#183c3a" stroke-width="6"/><circle cx="79" cy="111" r="5" fill="#fff"/><circle cx="119" cy="111" r="5" fill="#fff"/><path d="M80 130c10 11 27 11 37 0" fill="none" stroke="#fff" stroke-width="6" stroke-linecap="round"/></svg>
|
||||
|
After Width: | Height: | Size: 602 B |
1
assets/icon-512.svg
Normal file
|
|
@ -0,0 +1 @@
|
|||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 512 512"><rect width="512" height="512" rx="122" fill="#f5c95b"/><path d="M115 384c-14-53 19-96 69-107-27-26-18-64 18-83-7-34 19-70 57-70 48 0 73 43 55 81 46 13 65 58 41 91 46 18 64 58 44 97-27 59-250 59-284-9Z" fill="#75452f" stroke="#3f251e" stroke-width="16"/><path d="M176 139c47-54 162-58 203 16-54-18-123-16-187 16Z" fill="#28a9a1" stroke="#183c3a" stroke-width="16"/><circle cx="210" cy="296" r="14" fill="#fff"/><circle cx="317" cy="296" r="14" fill="#fff"/><path d="M214 347c26 29 72 29 99 0" fill="none" stroke="#fff" stroke-width="16" stroke-linecap="round"/></svg>
|
||||
|
After Width: | Height: | Size: 630 B |
9
assets/timmy.svg
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 220 220" role="img" aria-labelledby="title desc">
|
||||
<title id="title">Timmy the Talking Turd</title><desc id="desc">A smiling brown swirl wearing a teal cap</desc>
|
||||
<defs><linearGradient id="b" x1="0" y1="0" x2="1" y2="1"><stop stop-color="#9b5f3d"/><stop offset="1" stop-color="#5d3427"/></linearGradient></defs>
|
||||
<ellipse cx="110" cy="190" rx="72" ry="13" fill="#2d211b" opacity=".12"/>
|
||||
<path d="M52 165c-6-24 9-43 31-48-12-11-8-29 8-37-3-15 9-31 25-31 21 0 32 19 24 36 21 6 29 26 18 41 20 8 28 26 20 43-12 25-111 25-126-4Z" fill="url(#b)" stroke="#3f251e" stroke-width="6" stroke-linejoin="round"/>
|
||||
<path d="M78 54c20-24 72-26 89 7-23-8-54-7-82 7Z" fill="#28a9a1" stroke="#183c3a" stroke-width="6" stroke-linejoin="round"/><path d="M83 66c-4 3-13 5-25 1 8-10 18-15 28-14" fill="#28a9a1" stroke="#183c3a" stroke-width="6" stroke-linecap="round"/>
|
||||
<circle cx="91" cy="126" r="8" fill="#fff"/><circle cx="91" cy="128" r="4" fill="#27201c"/><circle cx="138" cy="126" r="8" fill="#fff"/><circle cx="138" cy="128" r="4" fill="#27201c"/>
|
||||
<path d="M92 150c11 13 31 13 43 0" fill="none" stroke="#fff" stroke-width="7" stroke-linecap="round"/><circle cx="72" cy="145" r="8" fill="#ed8e78" opacity=".65"/><circle cx="157" cy="145" r="8" fill="#ed8e78" opacity=".65"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 1.3 KiB |
23
citations.json
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
{
|
||||
"version": 1,
|
||||
"sources": [
|
||||
{
|
||||
"id": 1,
|
||||
"url": "https://www.continence.org.au/about-incontinence/bowel-incontinence/bristol-stool-chart",
|
||||
"title": "Bristol stool chart — Continence Health Australia",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"url": "https://www.nhs.uk/conditions/bleeding-from-the-bottom-rectal-bleeding",
|
||||
"title": "Bleeding from the bottom — NHS",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"url": "https://www.niddk.nih.gov/health-information/digestive-diseases/constipation/symptoms-causes",
|
||||
"title": "Symptoms & Causes of Constipation — NIDDK",
|
||||
"accessed": "2026-08-19"
|
||||
}
|
||||
]
|
||||
}
|
||||
18
index.html
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
<!doctype html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width,initial-scale=1,viewport-fit=cover">
|
||||
<meta name="theme-color" content="#f7f3ea">
|
||||
<meta name="description" content="A private, playful bowel diary that stays on your device.">
|
||||
<title>Timmy the Talking Turd</title>
|
||||
<link rel="manifest" href="/manifest.webmanifest">
|
||||
<link rel="icon" href="/assets/timmy.svg" type="image/svg+xml">
|
||||
<link rel="stylesheet" href="/styles.css">
|
||||
</head>
|
||||
<body>
|
||||
<div id="app" class="app-shell" aria-live="polite"></div>
|
||||
<noscript>Timmy needs JavaScript to keep your private diary on this device.</noscript>
|
||||
<script type="module" src="/app.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
13
manifest.webmanifest
Normal file
|
|
@ -0,0 +1,13 @@
|
|||
{
|
||||
"name": "Timmy the Talking Turd",
|
||||
"short_name": "Timmy",
|
||||
"description": "Your private, intelligent pooping pal.",
|
||||
"start_url": "/",
|
||||
"display": "standalone",
|
||||
"background_color": "#f7f3ea",
|
||||
"theme_color": "#f7f3ea",
|
||||
"icons": [
|
||||
{"src":"/assets/icon-192.svg","sizes":"192x192","type":"image/svg+xml","purpose":"any maskable"},
|
||||
{"src":"/assets/icon-512.svg","sizes":"512x512","type":"image/svg+xml","purpose":"any maskable"}
|
||||
]
|
||||
}
|
||||
62
package-lock.json
generated
Normal file
|
|
@ -0,0 +1,62 @@
|
|||
{
|
||||
"name": "timmy-talking-turd",
|
||||
"version": "0.1.0",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "timmy-talking-turd",
|
||||
"version": "0.1.0",
|
||||
"devDependencies": {
|
||||
"playwright": "^1.62.1"
|
||||
}
|
||||
},
|
||||
"node_modules/fsevents": {
|
||||
"version": "2.3.2",
|
||||
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.2.tgz",
|
||||
"integrity": "sha512-xiqMQR4xAeHTuB9uWm+fFRcIOgKBMiOBP+eXiyT7jsgVCq1bkVygt00oASowB7EdtpOHaaPgKt812P9ab+DDKA==",
|
||||
"dev": true,
|
||||
"hasInstallScript": true,
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
"darwin"
|
||||
],
|
||||
"engines": {
|
||||
"node": "^8.16.0 || ^10.6.0 || >=11.0.0"
|
||||
}
|
||||
},
|
||||
"node_modules/playwright": {
|
||||
"version": "1.62.1",
|
||||
"resolved": "https://registry.npmjs.org/playwright/-/playwright-1.62.1.tgz",
|
||||
"integrity": "sha512-0M+L3LAD8/nm554LOla9Ayx0j0tmFZ0FBcoQ7F1VuVHpM/XpiC8RcDzBQB8W5+hA8L22THxELzeF+2WcUzvcLg==",
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"dependencies": {
|
||||
"playwright-core": "1.62.1"
|
||||
},
|
||||
"bin": {
|
||||
"playwright": "cli.js"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=20"
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"fsevents": "2.3.2"
|
||||
}
|
||||
},
|
||||
"node_modules/playwright-core": {
|
||||
"version": "1.62.1",
|
||||
"resolved": "https://registry.npmjs.org/playwright-core/-/playwright-core-1.62.1.tgz",
|
||||
"integrity": "sha512-wPYSwEBJY9GHraISXqyqtx0na0LpO3XEX7jNDhntbex7tzUS7kLnZsOlFruFJB4Hi/rhDMjXGqHewDZ68nYZVw==",
|
||||
"dev": true,
|
||||
"license": "Apache-2.0",
|
||||
"bin": {
|
||||
"playwright-core": "cli.js"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=20"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
15
package.json
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
{
|
||||
"name": "timmy-talking-turd",
|
||||
"version": "0.1.0",
|
||||
"private": true,
|
||||
"type": "module",
|
||||
"scripts": {
|
||||
"test": "node --test tests/domain.test.js tests/analysis.test.js tests/vision-service.test.js tests/vision-config.test.js tests/training-ingest.test.js",
|
||||
"test:ui": "node tests/ui.acceptance.mjs",
|
||||
"test:photo": "node tests/photo-first.acceptance.mjs",
|
||||
"start": "node server.mjs"
|
||||
},
|
||||
"devDependencies": {
|
||||
"playwright": "^1.62.1"
|
||||
}
|
||||
}
|
||||
177
research/SELF-HOSTED-STOOL-VISION.md
Normal file
|
|
@ -0,0 +1,177 @@
|
|||
# Self-Hosted Stool Vision Decision — 2026-08-19
|
||||
|
||||
## Executive verdict
|
||||
|
||||
**Self-hosting removes the provider-shyness roadblock, but it does not remove the accuracy roadblock.** Timmy now has a working self-hosted profile, a local `llama.cpp` server path—using its OpenAI-compatible server capability[6]—provider readiness reporting, a consent-traceable training-image ingester, and a real-photo runtime receipt.
|
||||
|
||||
The recommended bootstrap model is **SmolVLM2-2.2B-Instruct Q4_K_M plus its Q8 vision projector**. The base model is Apache-2.0 and explicitly accepts image and text inputs.[11] The official GGUF repository provides the quantized main model and projector needed by `llama.cpp`.[10]
|
||||
|
||||
Do not ship this general-purpose VLM as the final Bristol classifier. On a real public CC-licensed Type 4 stool photograph, it did not refuse the image and emitted valid JSON, but it estimated Type 1 with confidence 0.5. Timmy's validator correctly abstained. That proves sovereign transport and content acceptance; it disproves product-grade classification accuracy.
|
||||
|
||||
## What was actually executed
|
||||
|
||||
### Host
|
||||
|
||||
- 4 virtual CPU cores at 2.0 GHz
|
||||
- 7.8 GiB RAM
|
||||
- no GPU
|
||||
- no swap
|
||||
|
||||
### Real stool input
|
||||
|
||||
The test image was the Wikimedia Commons file **Human Feces Bristol Stool Chart Type 4**, labeled Type 4 by its source and licensed CC BY-SA 3.0.[7] A derived test copy was resized, re-encoded, and stripped of EXIF before inference.
|
||||
|
||||
### SmolVLM-500M-Instruct Q8
|
||||
|
||||
The 500M model is Apache-2.0, is designed for image/text input, and its card says one-image inference can use 1.23 GB of GPU RAM.[1] The official GGUF files used in the spike were a 436,806,912-byte main model and a 108,783,360-byte Q8 projector.[2]
|
||||
|
||||
Observed on this CPU host:
|
||||
|
||||
- resident memory: **908,420 KiB**
|
||||
- structured request latency: **25.606 s**
|
||||
- detected stool: **yes**
|
||||
- Bristol type: **no usable value**
|
||||
- schema: malformed/incomplete
|
||||
- refusal: **none**
|
||||
|
||||
Verdict: useful only as a cheap stool/not-stool or quality-gate experiment; too weak for the product label.
|
||||
|
||||
### SmolVLM2-2.2B-Instruct Q4_K_M
|
||||
|
||||
Observed on this CPU host:
|
||||
|
||||
- model plus projector on disk: about **1.6 GiB**
|
||||
- resident memory: **3,362,048 KiB**
|
||||
- cold structured request: **89.259 s**
|
||||
- subsequent Timmy end-to-end request: **15.904 s**
|
||||
- detected stool: **yes**
|
||||
- predicted color: **brown**
|
||||
- predicted Bristol type: **1** (source label: Type 4)
|
||||
- confidence: **0.5**
|
||||
- image quality: **poor**
|
||||
- Timmy result: **abstained and requested user input**
|
||||
- refusal: **none**
|
||||
|
||||
Verdict: **adopt for bootstrap labeling experiments, reject as the final classifier.** The current VPS proves the path but lacks the latency and operational headroom for a smooth production service.
|
||||
|
||||
## Candidate decision
|
||||
|
||||
| Candidate | License observed | Current weight footprint | Decision |
|
||||
|---|---|---:|---|
|
||||
| SmolVLM2-2.2B-Instruct | Apache-2.0[11] | 1.6 GiB GGUF bundle used | **Bootstrap choice**; no refusal, structured output, inadequate Bristol accuracy |
|
||||
| SmolVLM-500M-Instruct | Apache-2.0[1] | 521 MiB GGUF bundle used | Quality/admission experiments only |
|
||||
| Moondream 2 | Apache-2.0, but its own card calls it the previous generation[5] | 3.85 GB BF16 model file observed | Do not start new work here |
|
||||
| Qwen2.5-VL-3B-Instruct | Qwen Research License[3] | 7.51 GB BF16 shards observed | **Reject for product default**: the license defines use as non-commercial research/evaluation and requires a separate license for commercial use.[4] |
|
||||
|
||||
## Production architecture
|
||||
|
||||
```text
|
||||
phone camera
|
||||
-> client resize + metadata removal
|
||||
-> Timmy ingress: decode, MIME/magic-byte check, size cap, request ID
|
||||
-> self-hosted visual gate: stool / not-stool / unusable
|
||||
-> specialist classifier: Type 1..7 probabilities + color probabilities
|
||||
-> calibration + abstention policy
|
||||
-> strict Timmy JSON schema
|
||||
-> user reviews and confirms
|
||||
-> local journal save
|
||||
-> separate, explicit opt-in training contribution
|
||||
```
|
||||
|
||||
### Model responsibilities
|
||||
|
||||
1. **General VLM:** bootstrap descriptions, reject non-stool images, flag blur/occlusion, and help reviewers. It never supplies the authoritative training label.
|
||||
2. **Human confirmation:** the user selects or corrects Bristol type and color. A clinical reviewer adjudicates the evaluation set and ambiguous examples.
|
||||
3. **Specialist discriminative model:** train a small image classifier for `not-stool`, `unknown`, and Types 1–7, with separate color and quality heads. This is the production inference model.
|
||||
4. **Timmy policy layer:** apply calibrated confidence thresholds, abstain, require confirmation, and keep symptoms entirely user-entered.
|
||||
|
||||
This structure matters because published stool-image work exists—including a long-term automated smart-toilet feasibility study[9]—but accessible training data is scarce. One 2024 pilot analyzed 151 images from only five hospitalized patients, and the paper says its data is not publicly available except for a possible limited request.[8] We therefore need our own consented dataset rather than pretending a generic VLM already solves the domain.
|
||||
|
||||
## Concrete consented-data pipeline
|
||||
|
||||
Implemented now: `scripts/ingest_training_photo.py`.
|
||||
|
||||
For every explicitly contributed image it:
|
||||
|
||||
- requires a consent version and human-review status;
|
||||
- decodes the image rather than trusting its extension;
|
||||
- applies EXIF orientation, then strips metadata;
|
||||
- resizes to a maximum 1200-pixel dimension;
|
||||
- hashes the derived bytes;
|
||||
- creates a salted pseudonymous subject key;
|
||||
- assigns train/validation/test by subject, not by image;
|
||||
- records Bristol type, broad color, quality, and review provenance in JSONL;
|
||||
- never writes the original source path into the manifest.
|
||||
|
||||
Keeping every image from one subject in one split prevents the model from looking better merely because near-duplicate photos from the same person leaked into training and test sets.
|
||||
|
||||
## Dataset ladder
|
||||
|
||||
### Phase A — 100-image reality check
|
||||
|
||||
- Collect opt-in images across all seven types and non-stool/poor-quality negatives.
|
||||
- Every label is user-confirmed; ambiguous items are clinician-reviewed.
|
||||
- Compare the VLM, a frozen image encoder plus linear head, and a small fine-tuned classifier.
|
||||
- Stop if the labels are too inconsistent to support the seven-way task.
|
||||
|
||||
### Phase B — minimum credible training corpus
|
||||
|
||||
Target at least:
|
||||
|
||||
- **250 independently reviewed examples per Bristol type**;
|
||||
- **200+ distinct contributors**;
|
||||
- hard negatives: empty toilets, toilet paper, urine-only, glare, blur, occlusion, diapers, and non-toilet contexts;
|
||||
- held-out contributors and devices;
|
||||
- intentionally varied lighting, bowl colors, water tint, distance, and camera quality.
|
||||
|
||||
These are engineering gates, not a guarantee of medical validity.
|
||||
|
||||
### Phase C — release gates
|
||||
|
||||
A build cannot prefill unless it demonstrates on held-out contributors:
|
||||
|
||||
- macro-F1 of at least **0.80** across Types 1–7;
|
||||
- at least **0.90 precision** among non-abstained suggestions;
|
||||
- expected calibration error at or below **0.05**;
|
||||
- valid schema on **100%** of responses;
|
||||
- no original-image persistence in transient inference;
|
||||
- no training use without an explicit contribution record;
|
||||
- zero automatic symptom, disease, or bleeding conclusions.
|
||||
|
||||
If confidence is below the calibrated threshold, Timmy asks the user. Coverage is allowed to fall; false certainty is not.
|
||||
|
||||
## Hardware gate
|
||||
|
||||
- **Current VPS:** accepted for a single-user engineering spike only. It ran the model, but 89-second cold latency, 3.36 GiB model RSS, four saturated CPU cores, no swap, and no GPU leave inadequate production margin.
|
||||
- **Bootstrap inference / fine-tuning:** use a separate NVIDIA GPU worker with **16 GB VRAM recommended**. Bind inference privately; Timmy remains the public API and release authority.
|
||||
- **Later production classifier:** after collecting data, export the specialist model to ONNX/TensorRT. A small classifier should be much faster and cheaper than keeping a VLM in the hot path, and may ultimately run on CPU or device.
|
||||
|
||||
## Shipped controls
|
||||
|
||||
- `TIMMY_VISION_PROFILE=selfhost`
|
||||
- loopback OpenAI-compatible default: `http://127.0.0.1:8080/v1`
|
||||
- self-hosted privacy status from `/api/vision-status`
|
||||
- upstream `/v1/models` readiness probe
|
||||
- API-key redaction from public status
|
||||
- 120-second bounded timeout for self-hosted inference
|
||||
- `scripts/run_selfhost_smolvlm.sh`
|
||||
- `scripts/ingest_training_photo.py`
|
||||
- fail-closed visual schema and manual fallback
|
||||
|
||||
## Decision
|
||||
|
||||
**Proceed with self-hosting, but treat SmolVLM2 as the bootstrap worker—not the product brain.** It solves “the provider refuses the shit picture.” The consented dataset plus specialist classifier solves “turn the shit picture into reliable concrete data.”
|
||||
|
||||
## Sources
|
||||
|
||||
[1] https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct — SmolVLM-500M-Instruct model card
|
||||
[2] https://huggingface.co/ggml-org/SmolVLM-500M-Instruct-GGUF — Official SmolVLM GGUF conversion
|
||||
[3] https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct — Qwen2.5-VL-3B-Instruct model card
|
||||
[4] https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE — Qwen research license
|
||||
[5] https://huggingface.co/vikhyatk/moondream2 — Moondream 2 model card
|
||||
[6] https://github.com/ggml-org/llama.cpp — llama.cpp repository
|
||||
[7] https://commons.wikimedia.org/wiki/File:Human_Feces_Bristol_Stool_Chart_Type_4.jpg — Wikimedia Commons: Human Feces Bristol Stool Chart Type 4
|
||||
[8] https://europepmc.org/articles/PMC11350077 — AI- and physician-interpreted stool image pilot study
|
||||
[9] https://europepmc.org/articles/PMC11686500 — Long-term automated stool monitoring smart toilet feasibility study
|
||||
[10] https://huggingface.co/ggml-org/SmolVLM2-2.2B-Instruct-GGUF — Official SmolVLM2 2.2B GGUF conversion
|
||||
[11] https://huggingface.co/HuggingFaceTB/SmolVLM2-2.2B-Instruct — SmolVLM2-2.2B-Instruct model card
|
||||
71
research/citations-ledger.json
Normal file
|
|
@ -0,0 +1,71 @@
|
|||
{
|
||||
"version": 1,
|
||||
"sources": [
|
||||
{
|
||||
"id": 1,
|
||||
"url": "https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct",
|
||||
"title": "SmolVLM-500M-Instruct model card",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"url": "https://huggingface.co/ggml-org/SmolVLM-500M-Instruct-GGUF",
|
||||
"title": "Official SmolVLM GGUF conversion",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"url": "https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct",
|
||||
"title": "Qwen2.5-VL-3B-Instruct model card",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"url": "https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE",
|
||||
"title": "Qwen research license",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"url": "https://huggingface.co/vikhyatk/moondream2",
|
||||
"title": "Moondream 2 model card",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"url": "https://github.com/ggml-org/llama.cpp",
|
||||
"title": "llama.cpp repository",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"url": "https://commons.wikimedia.org/wiki/File:Human_Feces_Bristol_Stool_Chart_Type_4.jpg",
|
||||
"title": "Wikimedia Commons: Human Feces Bristol Stool Chart Type 4",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"url": "https://europepmc.org/articles/PMC11350077",
|
||||
"title": "AI- and physician-interpreted stool image pilot study",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"url": "https://europepmc.org/articles/PMC11686500",
|
||||
"title": "Long-term automated stool monitoring smart toilet feasibility study",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 10,
|
||||
"url": "https://huggingface.co/ggml-org/SmolVLM2-2.2B-Instruct-GGUF",
|
||||
"title": "Official SmolVLM2 2.2B GGUF conversion",
|
||||
"accessed": "2026-08-19"
|
||||
},
|
||||
{
|
||||
"id": 11,
|
||||
"url": "https://huggingface.co/HuggingFaceTB/SmolVLM2-2.2B-Instruct",
|
||||
"title": "SmolVLM2-2.2B-Instruct model card",
|
||||
"accessed": "2026-08-19"
|
||||
}
|
||||
]
|
||||
}
|
||||
46
research/source-pages/PMC11350077.txt
Normal file
|
|
@ -0,0 +1,46 @@
|
|||
Address correspondence to: Gil Y. Melmed, MD, MS, Director of Inflammatory Bowel Disease Clinical Research, Cedars-Sinai Medical Center, 8720 Beverly Blvd, West Hollywood, CA, 90048, USA (310-423-4100, 310-423-0146, gil.melmed@cshs.org ).
|
||||
Stool characteristics are used as a measure of ulcerative colitis (UC) disease activity, but they have not been validated against objective inflammation. We aimed to determine whether stool characteristics measured by trained artificial intelligence (AI) and physicians correlate with inflammation in UC.
|
||||
Patients hospitalized with acute severe UC (ASUC) were asked to capture images of all bowel movements using a smartphone application (Dieta®). Validated AI was used to measure five stool characteristics including the Bristol stool scale. Additionally, four physicians scored each image for blood amount, mucus amount, and whether stool was in a toilet or commode. AI measurements and mean physician scores were rank-normalized and correlated with rank-normalized CRP values using mixed linear regression models. Mann–Whitney tests were used to compare median CRP values of images with and without mucus and with and without blood.
|
||||
We analyzed 151 stool images collected from 5 patients admitted with ASUC (mean age 42 years, 40% male). Overall, Bristol stool scale and fragmentation positively correlated with CRP, while stool consistency negatively correlated with CRP. The median CRP of images with mucus was higher than that of images without mucus.
|
||||
Smartphone application AI measurements of Bristol stool scale, stool consistency, and stool fragmentation significantly correlate with CRP values in hospitalized patients with ASUC. Additionally, median CRPs are higher when mucus is seen. Further training of smartphone-based AI algorithms to validate the association of stool characteristics with objective inflammation may yield a novel, noninvasive tool for UC disease monitoring.
|
||||
Keywords: Ulcerative Colitis, artificial intelligence, C-reactive protein, disease monitoring
|
||||
Received 2024 Jan 2; Collection date 2024 Jul.
|
||||
The management of acute, severe ulcerative colitis (ASUC) in the hospital warrants daily assessment and evaluation of disease activity to determine the appropriateness of continuing corticosteroids, initiating rescue therapy, or surgery. 1 Daily changes in patient-reported rectal bleeding, stool frequency, and form assist with clinical decision-making in the context of objective changes in inflammatory markers, but may be subject to interpretation and recall bias. Smart toilet technology may mitigate recall bias 2 ; however, this requires the installation of special technology into existing plumbing. We previously study that more easily accessible technology, smartphone-based artificial intelligence (AI), could determine stool characteristics with high accuracy and was superior to patient self-reporting. 3 We aimed to explore the clinical utility of AI applied to stool images acquired via smartphone application in patients hospitalized with ASUC.
|
||||
This prospective observational pilot study was reviewed and approved by the Cedars-Sinai Institutional Review Board. Informed consent was obtained prior to patient participation. Consecutive patients admitted to the hospital with ASUC captured images of each bowel movement using a smartphone application (Dieta®). The Dieta® mobile application was designed for patients to capture images of stool that are then classified by computer vision AI into multiple continuous data points for five different visual characteristics of stool. The AI classified stool image characteristics were Bristol stool scale (1 to 7), consistency (0, liquid to 100, solid), edge fuzziness (0, very clear to 100, very fuzzy), fragmentation (0, single piece to 100, many pieces), and volume (0, very small volume to 100, very large volume). Representative images of these characteristics can be found in Pimental et al. 3
|
||||
Stool images were also annotated by humans using the Dieta Stool Annotation Portal. This web application allows users to observe multiple views of each stool image and adjust brightness, contrast, saturation, and resolution. Using an annotation guide developed for this study, four physicians, including an inflammatory bowel disease specialist, labeled stool images for blood amount, mucus amount, and whether stool was in a toilet or bedside commode. Representative images of these characteristics can be found in Supplementary Figure 1 . Discrepancies were resolved by consensus.
|
||||
For each participant, patient demographics and relevant clinical information were recorded. Demographics included age, gender, race, and ethnicity. On admission, vitals signs, stool frequency, quantity of blood in stool, hemoglobin, and erythrocyte sedimentation rate were recorded to establish Truelove and Witts Severity Index for ulcerative colitis. Serum C-reactive protein (CRP) was measured on admission and daily while hospitalized. We additionally recorded fecal calprotectin, the presence of enteric infections, treatment received, and length of hospitalization.
|
||||
Each stool image obtained was associated with a serum CRP value obtained within 12 hours of the image. Image characteristics determined by AI and physician ratings were rank-normalized and correlated with rank-normalized CRP values using mixed linear regression models. We additionally used Mann–Whitney tests to compare median CRP values of images with the presence or absence of blood and/or mucus.
|
||||
Area Under the Receiver Operating Characteristic was calculated to evaluate the predictive performance of the linear mixed models used to correlate AI-assessed stool characteristics and CRP levels. CRP levels were dichotomized by median, tagging levels as below or equal and above median. Subsequently, predicted values were derived from each fitted linear mixed-effects model. The ROC curve was constructed utilizing the predicted values and the dichotomized CRP variable. The AUC value was calculated to provide a quantitative measure of the model’s discriminative ability in distinguishing between CRP levels below and above the median based on each AI-classified stool image characteristic.
|
||||
In total, 151 stool images were collected from 5 patients (mean age 42 years, 40% male). Detailed demographics and clinical information can be found in Table 1 . Each patient provided between 1 and 98 stool images, with a mean of 30.2 images per patient. On admission, patients had a mean CRP of 51.46 mg/L and a mean fecal calprotectin of 1684.8 μg/g. The baseline Truelove and Witts Severity Index was either moderate or severe. Patients were hospitalized for an average of 12.4 days. Campylobacter infection was detected in 1 patient and cytomegalovirus viral inclusions were present in 2 patients. 4 of the 5 patients received intravenous corticosteroids and infliximab. Fifty-three images were captured in a toilet; the remaining images were captured in a bedside commode.
|
||||
Demographics and clinical information.
|
||||
Serum CRP was positively correlated with AI-interpreted Bristol stool scale ( P = .026), stool consistency ( P = .047), and stool fragmentation ( P = .049), but not with stool volume or edge fuzziness. On subgroup analysis, only images obtained in a toilet maintained a statistically significant association ( Table 2 ). Additionally, the median CRP corresponding to images without blood was marginally higher than that of images with blood ( P = .07502; Figure 1A ) and the median CRP of images with mucus was significantly higher than that of images without mucus ( P = .01083; Figure 1B ).
|
||||
|
||||
P -values obtained by mixed linear regression models of CRP versus stool characteristics.
|
||||
|
||||
* Represents statistically significant P -value.
|
||||
Median serum CRP (mg/L) of stool images in which blood was or was not observed (A) and median serum CRP (mg/L) of stool images in which mucus was or was not observed (B).
|
||||
The AUC values for CRP with the Bristol scale, consistency, and fragmentation were 0.70, 0.691, and 0.70, respectively. Specifically focusing on images captured within a toilet environment, the AUC values for CRP in combination with the Bristol scale, consistency, and fragmentation were notably higher, measuring 0.81, 0.80, and 0.82, respectively ( Figure 2 ).
|
||||
Area under the receiver operating characteristic (AUROC) curves of the artificial intelligence (AI) model for prediction of CRP by the AI-predicted stool characteristics (Bristol, consistency, and fragmentation). Predictive performance of the AI model in the classification of CRP above or below the median. True positive rate of CRP classification performance as the x -axis and false positive rate as the y -axis. AUC ranges from ~0.7 to 0.8 when evaluating each model, which indicates a good rate of True positive CRP predictions.
|
||||
We found that AI classification of stool Bristol scale, consistency, and fragmentation obtained via smartphone application may correlate well with serum CRP in ASUC patients, with AUCs ranging from 0.6891 to 0.8211. The AI performed best using images of stools in toilets compared to images obtained in commodes likely due to the AI having been trained using images of stools in toilets alone. In the future, AI could also be trained to classify stool images for blood and mucus amounts. Notably, interpretation of the results of this study is limited by the small sample size and potentially confounding infections identified. Additionally, the AI utilized in this study was validated in a population of irritable bowel syndrome patients and therefore may not directly translate to the ASUC population. Large studies should be pursued to further validate the use of AI classification of stool images for use in ASUC disease monitoring.
|
||||
Jeff Liang, MD, Cedars-Sinai Medical Center, contributed to the investigation. Eden Sharabi, MD, Cedars-Sinai Medical Center, contributed to the investigation. Ryan Urbanowicz, PhD, Cedars-Sinai Medical Center, contributed to the formal analysis.
|
||||
Sarah Rotondo-Trivette,
|
||||
F. Widjaja Inflammatory Bowel Disease Institute, Department of Medicine, Cedars-Sinai, Los Angeles, CA, USA.
|
||||
Viankail Cedillo Castelan,
|
||||
F. Widjaja Inflammatory Bowel Disease Institute, Department of Medicine, Cedars-Sinai, Los Angeles, CA, USA.
|
||||
Kushagra Mathur,
|
||||
F. Widjaja Inflammatory Bowel Disease Institute, Department of Medicine, Cedars-Sinai, Los Angeles, CA, USA.
|
||||
Pauline Yasmeh,
|
||||
Department of Medicine, Olive View-UCLA Medical Center, Sylmar, CA, USA.
|
||||
Asaf Kraus,
|
||||
Dieta Health, Los Angeles, CA, USA.
|
||||
Addison Lynch,
|
||||
Dieta Health, Los Angeles, CA, USA.
|
||||
Dermot P B McGovern,
|
||||
F. Widjaja Inflammatory Bowel Disease Institute, Department of Medicine, Cedars-Sinai, Los Angeles, CA, USA.
|
||||
Gil Y Melmed,
|
||||
F. Widjaja Inflammatory Bowel Disease Institute, Department of Medicine, Cedars-Sinai, Los Angeles, CA, USA.
|
||||
Sarah Rotondo-Trivette contributed to conceptualization, methodology, formal analysis, investigation, resources, data curation, project administration, and writing. Viankail Cedillo Castelan contributed to formal analysis, data curation, and visualization. Kushagra Mathur contributed to the investigation, data curation, and review/editing. Pauline Yasmeh contributed to the investigation, data curation, and review/editing. Asaf Kraus contributed to methodology, software, resources, and writing. Addison Lynch contributed to software and resources. Dermot P.B. McGovern contributed to conceptualization, methodology, and review/editing. Gil Y. Melmed contributed to conceptualization, methodology, formal analysis, investigation, resources, writing, supervision, and project administration. All authors have approved the final draft submitted.
|
||||
None.
|
||||
Sarah Rotondo-Trivette, Viankail Cedillo Castelan, Kushagra Mathur, and Pauline Yasmeh have no conflicts of interest to disclose. Asaf Kraus discloses their position as shareholder, board member, and employee of Dieta Health. Addison Lynch discloses position as an employee of Dieta Health. Dermot P.B. McGovern discloses position as consultant to MERCK, Palisade Bio, Prometheus Biosciences, Prometheus Labs, and Takeda. Gil Y Melmed discloses position as consultant to Abbvie, Arena, Boehringer-Ingelheim, Bristol-Myers Squibb, Celgene, Dieta Health, Entasis, Ferring, Fresenius Kabi, Janssen, Medtronic, Oshi Health, Pfizer, Takeda, Shionogi, Samsung Bioepis, and Viatris.
|
||||
Data is not publicly available but a limited dataset could be provided upon reasonable request to the corresponding author.
|
||||
Data is not publicly available but a limited dataset could be provided upon reasonable request to the corresponding author.
|
||||
126
research/source-pages/llama.md
Normal file
|
|
@ -0,0 +1,126 @@
|
|||
# llama.cpp
|
||||
|
||||

|
||||
|
||||
<div align="center">
|
||||
|
||||
<b>LLM inference in C/C++</b>
|
||||
|
||||
[](https://opensource.org/licenses/MIT)
|
||||
[](https://github.com/ggml-org/llama.cpp/releases?q=tag:v0)
|
||||
[](https://github.com/ggml-org/llama.cpp/releases)
|
||||
[](https://github.com/ggml-org/llama.cpp/actions/workflows/server.yml)
|
||||
[](https://github.com/ggml-org/llama.cpp/actions/workflows/docker.yml)
|
||||
[](https://github.com/ggml-org/llama.cpp/actions/workflows/winget.yml)
|
||||
|
||||
[manifesto](https://github.com/ggml-org/llama.cpp/discussions/205) / [ggml](https://github.com/ggml-org/ggml) / [ops](https://github.com/ggml-org/llama.cpp/blob/master/docs/ops.md) / [maintainer PRs](https://github.com/ggml-org/llama.cpp/issues?q=is%3Apr%20is%3Aopen%20draft%3AFalse%20(author%3Argerganov%20OR%20author%3AKitaitiMakoto%20OR%20author%3Adanbev%20OR%20author%3Aaldehir%20OR%20author%3Amax-krasnyansky%20OR%20author%3ACISC%20OR%20author%3Aggerganov%20OR%20author%3Aam17an%20OR%20author%3Abartowski1182%20OR%20author%3Ahipudding%20OR%20author%3AServeurpersoCom%20OR%20author%3Apwilkin%20OR%20author%3Areeselevine%20OR%20author%3Angxson%20OR%20author%3Ajeffbolznv%20OR%20author%3A0cc4m%20OR%20author%3Aangt%20OR%20author%3AIMbackK%20OR%20author%3Aarthw%20OR%20author%3AJohannesGaessler%20OR%20author%3AORippler%20OR%20author%3Aruixiang63%20OR%20author%3Axctan%20OR%20author%3Aallozaur%20OR%20author%3Ayomaytk%20OR%20author%3Aaendk%20OR%20author%3Agaugarg-nv%20OR%20author%3Ataronaeo%20OR%20author%3Aforforever73%20OR%20author%3Alhez%20OR%20author%3Anetrunnereve%20OR%20author%3Afairydreaming)%20sort%3Aupdated-desc) / [compile times](https://github.com/ggml-org/llama.cpp-dev/blob/master/README-compile-times.md) / [lib llama API](https://github.com/ggml-org/llama.cpp/issues/9289) / [llama-server REST API](https://github.com/ggml-org/llama.cpp/issues/9291)
|
||||
|
||||
</div>
|
||||
|
||||
## Quick start
|
||||
|
||||
A few options to get `llama.cpp` installed on your machine:
|
||||
|
||||
- Visit https://llama.app and follow the instructions
|
||||
- Run with Docker - see our [Docker documentation](docs/docker.md)
|
||||
- Download pre-built binaries from the [releases page](https://github.com/ggml-org/llama.cpp/releases)
|
||||
- Build from source by cloning this repository - check out [our build guide](docs/build.md)
|
||||
|
||||
Once installed:
|
||||
|
||||
```sh
|
||||
# Download and run a model directly from Hugging Face
|
||||
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
|
||||
|
||||
# Launch OpenAI-compatible API server
|
||||
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
||||
```
|
||||
|
||||
<table align="center">
|
||||
<tr>
|
||||
<td align="center" width=50%>
|
||||
<img width="1310" height="888" alt="VLM session with `llama cli`" src="https://github.com/user-attachments/assets/88726b48-1713-48aa-a525-95a02e78afc4" />
|
||||
<i>VLM session with <b>llama cli</b></i>
|
||||
</td>
|
||||
<td align="center">
|
||||
<img width="1392" height="958" alt="Built-in web UI against `llama serve` running Qwen 3.6" src="https://github.com/user-attachments/assets/b402f972-2e32-4def-8771-8d849f08cf2e" />
|
||||
<i>Built-in web UI against <b>llama serve</b></i>
|
||||
</td>
|
||||
</tr>
|
||||
<table>
|
||||
|
||||
## Description
|
||||
|
||||
The main goal of `llama.cpp` is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
|
||||
a wide range of hardware - locally and in the cloud.
|
||||
|
||||
- Plain C/C++ implementation without any dependencies
|
||||
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
|
||||
- AVX, AVX2, AVX512 and AMX support for x86 architectures
|
||||
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
|
||||
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
|
||||
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
|
||||
- Vulkan and SYCL backend support
|
||||
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
|
||||
|
||||
The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-org/ggml) library.
|
||||
|
||||
## Supported backends
|
||||
|
||||
| Backend | Target devices |
|
||||
| --- | --- |
|
||||
| [BLAS](docs/build.md#blas-build) | All |
|
||||
| [BLIS](docs/backend/BLIS.md) | All |
|
||||
| [CANN](docs/build.md#cann) | Ascend NPU |
|
||||
| [CUDA](docs/build.md#cuda) | Nvidia GPU |
|
||||
| [HIP](docs/build.md#hip) | AMD GPU |
|
||||
| [Hexagon [In Progress]](docs/backend/snapdragon/README.md) | Snapdragon |
|
||||
| [IBM zDNN](docs/backend/zDNN.md) | IBM Z & LinuxONE |
|
||||
| [MUSA](docs/build.md#musa) | Moore Threads GPU |
|
||||
| [Metal](docs/build.md#metal-build) | Apple Silicon |
|
||||
| [OpenCL](docs/backend/OPENCL.md) | Adreno GPU |
|
||||
| [OpenVINO [In Progress]](docs/backend/OPENVINO.md) | Intel CPUs, GPUs, and NPUs |
|
||||
| [RPC](https://github.com/ggml-org/llama.cpp/tree/master/tools/rpc) | All |
|
||||
| [SYCL](docs/backend/SYCL.md) | Intel GPU |
|
||||
| [VirtGPU](docs/backend/VirtGPU.md) | VirtGPU APIR |
|
||||
| [Vulkan](docs/build.md#vulkan) | GPU |
|
||||
| [WebGPU](docs/build.md#webgpu) | All |
|
||||
| [ZenDNN](docs/build.md#zendnn) | AMD CPU |
|
||||
|
||||
## Documentation
|
||||
|
||||
#### Tools
|
||||
|
||||
- [cli](tools/cli/README.md)
|
||||
- [completion](tools/completion/README.md)
|
||||
- [server](tools/server/README.md)
|
||||
- [GBNF grammars](grammars/README.md)
|
||||
|
||||
#### Development
|
||||
|
||||
- [How to build](docs/build.md)
|
||||
- [Running on Docker](docs/docker.md)
|
||||
- [Build on Android](docs/android.md)
|
||||
- [Multi-GPU usage](docs/multi-gpu.md)
|
||||
- [Performance troubleshooting](docs/development/token_generation_performance_tips.md)
|
||||
- [GGML tips & tricks](https://github.com/ggml-org/llama.cpp/wiki/GGML-Tips-&-Tricks)
|
||||
- [XCFramework](docs/xcframework.md)
|
||||
- [Completions](docs/completions.md)
|
||||
- [Models](docs/models.md)
|
||||
- [Release process](docs/release.md)
|
||||
|
||||
## Contributing
|
||||
|
||||
- Contributors can open PRs
|
||||
- Collaborators will be invited based on contributions
|
||||
- Maintainers can push to branches in the `llama.cpp` repo and merge PRs into the `master` branch
|
||||
- Any help with managing issues, PRs and projects is very appreciated!
|
||||
- Read the [CONTRIBUTING.md](CONTRIBUTING.md) for more information
|
||||
|
||||
## Acknowledgements
|
||||
|
||||
- [yhirose/cpp-httplib](https://github.com/yhirose/cpp-httplib) - Single-header HTTP server, used by `llama-server` - MIT license
|
||||
- [nothings/stb](https://github.com/nothings/stb) - Single-header image format decoder, used by multimodal subsystem - Public domain
|
||||
- [nlohmann/json](https://github.com/nlohmann/json) - Single-header JSON library, used by various tools/examples - MIT License
|
||||
- [mackron/miniaudio](https://github.com/mackron/miniaudio) - Single-header audio format decoder, used by multimodal subsystem - Public domain
|
||||
- [sheredom/subprocess.h](https://github.com/sheredom/subprocess.h) - Single-header process launching solution for C and C++ - Public domain
|
||||
88
research/source-pages/moondream.md
Normal file
|
|
@ -0,0 +1,88 @@
|
|||
---
|
||||
license: apache-2.0
|
||||
pipeline_tag: image-text-to-text
|
||||
new_version: moondream/moondream3-preview
|
||||
---
|
||||
|
||||
⚠️ This repository contains the latest version of Moondream 2, our previous generation model. The latest version of Moondream is [Moondream 3 (Preview)](https://huggingface.co/moondream/moondream3-preview).
|
||||
|
||||
---
|
||||
|
||||
Moondream is a small vision language model designed to run efficiently everywhere.
|
||||
|
||||
[Website](https://moondream.ai/) / [Demo](https://moondream.ai/playground) / [GitHub](https://github.com/vikhyat/moondream)
|
||||
|
||||
This repository contains the latest (**2025-06-21**) release of Moondream 2, as well as [historical releases](https://huggingface.co/vikhyatk/moondream2/blob/main/versions.txt). The model is updated frequently, so we recommend specifying a revision as shown below if you're using it in a production application.
|
||||
|
||||
|
||||
### Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
from PIL import Image
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
"vikhyatk/moondream2",
|
||||
revision="2025-06-21",
|
||||
trust_remote_code=True,
|
||||
device_map={"": "cuda"} # ...or 'mps', on Apple Silicon
|
||||
)
|
||||
|
||||
# Captioning
|
||||
print("Short caption:")
|
||||
print(model.caption(image, length="short")["caption"])
|
||||
|
||||
print("\nNormal caption:")
|
||||
for t in model.caption(image, length="normal", stream=True)["caption"]:
|
||||
# Streaming generation example, supported for caption() and detect()
|
||||
print(t, end="", flush=True)
|
||||
print(model.caption(image, length="normal"))
|
||||
|
||||
# Visual Querying
|
||||
print("\nVisual query: 'How many people are in the image?'")
|
||||
print(model.query(image, "How many people are in the image?")["answer"])
|
||||
|
||||
# Object Detection
|
||||
print("\nObject detection: 'face'")
|
||||
objects = model.detect(image, "face")["objects"]
|
||||
print(f"Found {len(objects)} face(s)")
|
||||
|
||||
# Pointing
|
||||
print("\nPointing: 'person'")
|
||||
points = model.point(image, "person")["points"]
|
||||
print(f"Found {len(points)} person(s)")
|
||||
```
|
||||
|
||||
### Changelog
|
||||
|
||||
**2025-06-21** ([full release notes](https://moondream.ai/blog/moondream-2025-06-21-release))
|
||||
|
||||
* **Grounded Reasoning**
|
||||
Introduces a new step-by-step reasoning mode that explicitly grounds reasoning in spatial positions within the image before answering, leading to more precise visual interpretation (e.g., chart median calculations, accurate counting). Enable with `reasoning=True` in the `query` skill to trade off speed vs. accuracy.
|
||||
* **Sharper Object Detection**
|
||||
Uses reinforcement learning on higher-quality bounding-box annotations to reduce object clumping and improve fine-grained detections (e.g., distinguishing “blue bottle” vs. “bottle”).
|
||||
* **Faster Text Generation**
|
||||
Yields 20–40 % faster response generation via a new “superword” tokenizer and lightweight tokenizer transfer hypernetwork, which reduces the number of tokens emitted without loss in accuracy and eases future multilingual extensions.
|
||||
* **Improved UI Understanding**
|
||||
Boosts ScreenSpot (UI element localization) performance from an F1\@0.5 of 60.3 to 80.4, making Moondream more effective for UI-focused applications.
|
||||
* **Reinforcement Learning Enhancements**
|
||||
RL fine-tuning applied across 55 vision-language tasks to reinforce grounded reasoning and detection capabilities, with a roadmap to expand to \~120 tasks in the next update.
|
||||
|
||||
**2025-04-15** ([full release notes](https://moondream.ai/blog/moondream-2025-04-14-release))
|
||||
|
||||
1. Improved chart understanding (ChartQA up from 74.8 to 77.5, 82.2 with PoT)
|
||||
2. Added temperature and nucleus sampling to reduce repetitive outputs
|
||||
3. Better OCR for documents and tables (prompt with “Transcribe the text” or “Transcribe the text in natural reading order”)
|
||||
4. Object detection supports document layout detection (figure, formula, text, etc)
|
||||
5. UI understanding (ScreenSpot F1\@0.5 up from 53.3 to 60.3)
|
||||
6. Improved text understanding (DocVQA up from 76.5 to 79.3, TextVQA up from 74.6 to 76.3)
|
||||
|
||||
**2025-03-27** ([full release notes](https://moondream.ai/blog/moondream-2025-03-27-release))
|
||||
|
||||
1. Added support for long-form captioning
|
||||
2. Open vocabulary image tagging
|
||||
3. Improved counting accuracy (e.g. CountBenchQA increased from 80 to 86.4)
|
||||
4. Improved text understanding (e.g. OCRBench increased from 58.3 to 61.2)
|
||||
5. Improved object detection, especially for small objects (e.g. COCO up from 30.5 to 51.2)
|
||||
6. Fixed token streaming bug affecting multi-byte unicode characters
|
||||
7. gpt-fast style `compile()` now supported in HF Transformers implementation
|
||||
525
research/source-pages/qwen.md
Normal file
|
|
@ -0,0 +1,525 @@
|
|||
|
||||
---
|
||||
license_name: qwen-research
|
||||
license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE
|
||||
language:
|
||||
- en
|
||||
pipeline_tag: image-text-to-text
|
||||
tags:
|
||||
- multimodal
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# Qwen2.5-VL-3B-Instruct
|
||||
<a href="https://chat.qwenlm.ai/" target="_blank" style="margin: 2px;">
|
||||
<img alt="Chat" src="https://img.shields.io/badge/%F0%9F%92%9C%EF%B8%8F%20Qwen%20Chat%20-536af5" style="display: inline-block; vertical-align: middle;"/>
|
||||
</a>
|
||||
|
||||
## Introduction
|
||||
|
||||
In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.
|
||||
|
||||
#### Key Enhancements:
|
||||
* **Understand things visually**: Qwen2.5-VL is not only proficient in recognizing common objects such as flowers, birds, fish, and insects, but it is highly capable of analyzing texts, charts, icons, graphics, and layouts within images.
|
||||
|
||||
* **Being agentic**: Qwen2.5-VL directly plays as a visual agent that can reason and dynamically direct tools, which is capable of computer use and phone use.
|
||||
|
||||
* **Understanding long videos and capturing events**: Qwen2.5-VL can comprehend videos of over 1 hour, and this time it has a new ability of cpaturing event by pinpointing the relevant video segments.
|
||||
|
||||
* **Capable of visual localization in different formats**: Qwen2.5-VL can accurately localize objects in an image by generating bounding boxes or points, and it can provide stable JSON outputs for coordinates and attributes.
|
||||
|
||||
* **Generating structured outputs**: for data like scans of invoices, forms, tables, etc. Qwen2.5-VL supports structured outputs of their contents, benefiting usages in finance, commerce, etc.
|
||||
|
||||
|
||||
#### Model Architecture Updates:
|
||||
|
||||
* **Dynamic Resolution and Frame Rate Training for Video Understanding**:
|
||||
|
||||
We extend dynamic resolution to the temporal dimension by adopting dynamic FPS sampling, enabling the model to comprehend videos at various sampling rates. Accordingly, we update mRoPE in the time dimension with IDs and absolute time alignment, enabling the model to learn temporal sequence and speed, and ultimately acquire the ability to pinpoint specific moments.
|
||||
|
||||
<p align="center">
|
||||
<img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5-VL/qwen2.5vl_arc.jpeg" width="80%"/>
|
||||
<p>
|
||||
|
||||
|
||||
* **Streamlined and Efficient Vision Encoder**
|
||||
|
||||
We enhance both training and inference speeds by strategically implementing window attention into the ViT. The ViT architecture is further optimized with SwiGLU and RMSNorm, aligning it with the structure of the Qwen2.5 LLM.
|
||||
|
||||
|
||||
We have three models with 3, 7 and 72 billion parameters. This repo contains the instruction-tuned 3B Qwen2.5-VL model. For more information, visit our [Blog](https://qwenlm.github.io/blog/qwen2.5-vl/) and [GitHub](https://github.com/QwenLM/Qwen2.5-VL).
|
||||
|
||||
|
||||
|
||||
## Evaluation
|
||||
|
||||
### Image benchmark
|
||||
|
||||
| Benchmark | InternVL2.5-4B |Qwen2-VL-7B |Qwen2.5-VL-3B |
|
||||
| :--- | :---: | :---: | :---: |
|
||||
| MMMU<sub>val</sub> | 52.3 | 54.1 | 53.1|
|
||||
| MMMU-Pro<sub>val</sub> | **32.7** | 30.5 | 31.6|
|
||||
| AI2D<sub>test</sub> | 81.4 | **83.0** | 81.5 |
|
||||
| DocVQA<sub>test</sub> | 91.6 | 94.5 | **93.9** |
|
||||
| InfoVQA<sub>test</sub> | 72.1 | 76.5 | **77.1** |
|
||||
| TextVQA<sub>val</sub> | 76.8 | **84.3** | 79.3|
|
||||
| MMBench-V1.1<sub>test</sub> | 79.3 | **80.7** | 77.6 |
|
||||
| MMStar | 58.3 | **60.7** | 55.9 |
|
||||
| MathVista<sub>testmini</sub> | 60.5 | 58.2 | **62.3** |
|
||||
| MathVision<sub>full</sub> | 20.9 | 16.3 | **21.2** |
|
||||
|
||||
|
||||
### Video benchmark
|
||||
| Benchmark | InternVL2.5-4B | Qwen2-VL-7B | Qwen2.5-VL-3B |
|
||||
| :--- | :---: | :---: | :---: |
|
||||
| MVBench | 71.6 | 67.0 | 67.0 |
|
||||
| VideoMME | 63.6/62.3 | 69.0/63.3 | 67.6/61.5 |
|
||||
| MLVU | 48.3 | - | 68.2 |
|
||||
| LVBench | - | - | 43.3 |
|
||||
| MMBench-Video | 1.73 | 1.44 | 1.63 |
|
||||
| EgoSchema | - | - | 64.8 |
|
||||
| PerceptionTest | - | - | 66.9 |
|
||||
| TempCompass | - | - | 64.4 |
|
||||
| LongVideoBench | 55.2 | 55.6 | 54.2 |
|
||||
| CharadesSTA/mIoU | - | - | 38.8 |
|
||||
|
||||
|
||||
### Agent benchmark
|
||||
| Benchmarks | Qwen2.5-VL-3B |
|
||||
|-------------------------|---------------|
|
||||
| ScreenSpot | 55.5 |
|
||||
| ScreenSpot Pro | 23.9 |
|
||||
| AITZ_EM | 76.9 |
|
||||
| Android Control High_EM | 63.7 |
|
||||
| Android Control Low_EM | 22.2 |
|
||||
| AndroidWorld_SR | 90.8 |
|
||||
| MobileMiniWob++_SR | 67.9 |
|
||||
|
||||
## Requirements
|
||||
The code of Qwen2.5-VL has been in the latest Hugging face transformers and we advise you to build from source with command:
|
||||
```
|
||||
pip install git+https://github.com/huggingface/transformers accelerate
|
||||
```
|
||||
or you might encounter the following error:
|
||||
```
|
||||
KeyError: 'qwen2_5_vl'
|
||||
```
|
||||
|
||||
|
||||
## Quickstart
|
||||
|
||||
Below, we provide simple examples to show how to use Qwen2.5-VL with 🤖 ModelScope and 🤗 Transformers.
|
||||
|
||||
The code of Qwen2.5-VL has been in the latest Hugging face transformers and we advise you to build from source with command:
|
||||
```
|
||||
pip install git+https://github.com/huggingface/transformers accelerate
|
||||
```
|
||||
or you might encounter the following error:
|
||||
```
|
||||
KeyError: 'qwen2_5_vl'
|
||||
```
|
||||
|
||||
|
||||
We offer a toolkit to help you handle various types of visual input more conveniently, as if you were using an API. This includes base64, URLs, and interleaved images and videos. You can install it using the following command:
|
||||
|
||||
```bash
|
||||
# It's highly recommanded to use `[decord]` feature for faster video loading.
|
||||
pip install qwen-vl-utils[decord]==0.0.8
|
||||
```
|
||||
|
||||
If you are not using Linux, you might not be able to install `decord` from PyPI. In that case, you can use `pip install qwen-vl-utils` which will fall back to using torchvision for video processing. However, you can still [install decord from source](https://github.com/dmlc/decord?tab=readme-ov-file#install-from-source) to get decord used when loading video.
|
||||
|
||||
### Using 🤗 Transformers to Chat
|
||||
|
||||
Here we show a code snippet to show you how to use the chat model with `transformers` and `qwen_vl_utils`:
|
||||
|
||||
```python
|
||||
from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
|
||||
from qwen_vl_utils import process_vision_info
|
||||
|
||||
# default: Load the model on the available device(s)
|
||||
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
|
||||
"Qwen/Qwen2.5-VL-3B-Instruct", torch_dtype="auto", device_map="auto"
|
||||
)
|
||||
|
||||
# We recommend enabling flash_attention_2 for better acceleration and memory saving, especially in multi-image and video scenarios.
|
||||
# model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
|
||||
# "Qwen/Qwen2.5-VL-3B-Instruct",
|
||||
# torch_dtype=torch.bfloat16,
|
||||
# attn_implementation="flash_attention_2",
|
||||
# device_map="auto",
|
||||
# )
|
||||
|
||||
# default processer
|
||||
processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct")
|
||||
|
||||
# The default range for the number of visual tokens per image in the model is 4-16384.
|
||||
# You can set min_pixels and max_pixels according to your needs, such as a token range of 256-1280, to balance performance and cost.
|
||||
# min_pixels = 256*28*28
|
||||
# max_pixels = 1280*28*28
|
||||
# processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels)
|
||||
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image",
|
||||
"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
|
||||
},
|
||||
{"type": "text", "text": "Describe this image."},
|
||||
],
|
||||
}
|
||||
]
|
||||
|
||||
# Preparation for inference
|
||||
text = processor.apply_chat_template(
|
||||
messages, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
image_inputs, video_inputs = process_vision_info(messages)
|
||||
inputs = processor(
|
||||
text=[text],
|
||||
images=image_inputs,
|
||||
videos=video_inputs,
|
||||
padding=True,
|
||||
return_tensors="pt",
|
||||
)
|
||||
inputs = inputs.to("cuda")
|
||||
|
||||
# Inference: Generation of the output
|
||||
generated_ids = model.generate(**inputs, max_new_tokens=128)
|
||||
generated_ids_trimmed = [
|
||||
out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
|
||||
]
|
||||
output_text = processor.batch_decode(
|
||||
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
|
||||
)
|
||||
print(output_text)
|
||||
```
|
||||
<details>
|
||||
<summary>Multi image inference</summary>
|
||||
|
||||
```python
|
||||
# Messages containing multiple images and a text query
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "image": "file:///path/to/image1.jpg"},
|
||||
{"type": "image", "image": "file:///path/to/image2.jpg"},
|
||||
{"type": "text", "text": "Identify the similarities between these images."},
|
||||
],
|
||||
}
|
||||
]
|
||||
|
||||
# Preparation for inference
|
||||
text = processor.apply_chat_template(
|
||||
messages, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
image_inputs, video_inputs = process_vision_info(messages)
|
||||
inputs = processor(
|
||||
text=[text],
|
||||
images=image_inputs,
|
||||
videos=video_inputs,
|
||||
padding=True,
|
||||
return_tensors="pt",
|
||||
)
|
||||
inputs = inputs.to("cuda")
|
||||
|
||||
# Inference
|
||||
generated_ids = model.generate(**inputs, max_new_tokens=128)
|
||||
generated_ids_trimmed = [
|
||||
out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
|
||||
]
|
||||
output_text = processor.batch_decode(
|
||||
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
|
||||
)
|
||||
print(output_text)
|
||||
```
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary>Video inference</summary>
|
||||
|
||||
```python
|
||||
# Messages containing a images list as a video and a text query
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "video",
|
||||
"video": [
|
||||
"file:///path/to/frame1.jpg",
|
||||
"file:///path/to/frame2.jpg",
|
||||
"file:///path/to/frame3.jpg",
|
||||
"file:///path/to/frame4.jpg",
|
||||
],
|
||||
},
|
||||
{"type": "text", "text": "Describe this video."},
|
||||
],
|
||||
}
|
||||
]
|
||||
|
||||
# Messages containing a local video path and a text query
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "video",
|
||||
"video": "file:///path/to/video1.mp4",
|
||||
"max_pixels": 360 * 420,
|
||||
"fps": 1.0,
|
||||
},
|
||||
{"type": "text", "text": "Describe this video."},
|
||||
],
|
||||
}
|
||||
]
|
||||
|
||||
# Messages containing a video url and a text query
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "video",
|
||||
"video": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2-VL/space_woaudio.mp4",
|
||||
},
|
||||
{"type": "text", "text": "Describe this video."},
|
||||
],
|
||||
}
|
||||
]
|
||||
|
||||
#In Qwen 2.5 VL, frame rate information is also input into the model to align with absolute time.
|
||||
# Preparation for inference
|
||||
text = processor.apply_chat_template(
|
||||
messages, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
image_inputs, video_inputs, video_kwargs = process_vision_info(messages, return_video_kwargs=True)
|
||||
inputs = processor(
|
||||
text=[text],
|
||||
images=image_inputs,
|
||||
videos=video_inputs,
|
||||
fps=fps,
|
||||
padding=True,
|
||||
return_tensors="pt",
|
||||
**video_kwargs,
|
||||
)
|
||||
inputs = inputs.to("cuda")
|
||||
|
||||
# Inference
|
||||
generated_ids = model.generate(**inputs, max_new_tokens=128)
|
||||
generated_ids_trimmed = [
|
||||
out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
|
||||
]
|
||||
output_text = processor.batch_decode(
|
||||
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
|
||||
)
|
||||
print(output_text)
|
||||
```
|
||||
|
||||
Video URL compatibility largely depends on the third-party library version. The details are in the table below. change the backend by `FORCE_QWENVL_VIDEO_READER=torchvision` or `FORCE_QWENVL_VIDEO_READER=decord` if you prefer not to use the default one.
|
||||
|
||||
| Backend | HTTP | HTTPS |
|
||||
|-------------|------|-------|
|
||||
| torchvision >= 0.19.0 | ✅ | ✅ |
|
||||
| torchvision < 0.19.0 | ❌ | ❌ |
|
||||
| decord | ✅ | ❌ |
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary>Batch inference</summary>
|
||||
|
||||
```python
|
||||
# Sample messages for batch inference
|
||||
messages1 = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "image": "file:///path/to/image1.jpg"},
|
||||
{"type": "image", "image": "file:///path/to/image2.jpg"},
|
||||
{"type": "text", "text": "What are the common elements in these pictures?"},
|
||||
],
|
||||
}
|
||||
]
|
||||
messages2 = [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": "Who are you?"},
|
||||
]
|
||||
# Combine messages for batch processing
|
||||
messages = [messages1, messages2]
|
||||
|
||||
# Preparation for batch inference
|
||||
texts = [
|
||||
processor.apply_chat_template(msg, tokenize=False, add_generation_prompt=True)
|
||||
for msg in messages
|
||||
]
|
||||
image_inputs, video_inputs = process_vision_info(messages)
|
||||
inputs = processor(
|
||||
text=texts,
|
||||
images=image_inputs,
|
||||
videos=video_inputs,
|
||||
padding=True,
|
||||
return_tensors="pt",
|
||||
)
|
||||
inputs = inputs.to("cuda")
|
||||
|
||||
# Batch Inference
|
||||
generated_ids = model.generate(**inputs, max_new_tokens=128)
|
||||
generated_ids_trimmed = [
|
||||
out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
|
||||
]
|
||||
output_texts = processor.batch_decode(
|
||||
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
|
||||
)
|
||||
print(output_texts)
|
||||
```
|
||||
</details>
|
||||
|
||||
### 🤖 ModelScope
|
||||
We strongly advise users especially those in mainland China to use ModelScope. `snapshot_download` can help you solve issues concerning downloading checkpoints.
|
||||
|
||||
|
||||
### More Usage Tips
|
||||
|
||||
For input images, we support local files, base64, and URLs. For videos, we currently only support local files.
|
||||
|
||||
```python
|
||||
# You can directly insert a local file path, a URL, or a base64-encoded image into the position where you want in the text.
|
||||
## Local file path
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "image": "file:///path/to/your/image.jpg"},
|
||||
{"type": "text", "text": "Describe this image."},
|
||||
],
|
||||
}
|
||||
]
|
||||
## Image URL
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "image": "http://path/to/your/image.jpg"},
|
||||
{"type": "text", "text": "Describe this image."},
|
||||
],
|
||||
}
|
||||
]
|
||||
## Base64 encoded image
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "image": "data:image;base64,/9j/..."},
|
||||
{"type": "text", "text": "Describe this image."},
|
||||
],
|
||||
}
|
||||
]
|
||||
```
|
||||
#### Image Resolution for performance boost
|
||||
|
||||
The model supports a wide range of resolution inputs. By default, it uses the native resolution for input, but higher resolutions can enhance performance at the cost of more computation. Users can set the minimum and maximum number of pixels to achieve an optimal configuration for their needs, such as a token count range of 256-1280, to balance speed and memory usage.
|
||||
|
||||
```python
|
||||
min_pixels = 256 * 28 * 28
|
||||
max_pixels = 1280 * 28 * 28
|
||||
processor = AutoProcessor.from_pretrained(
|
||||
"Qwen/Qwen2.5-VL-3B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels
|
||||
)
|
||||
```
|
||||
|
||||
Besides, We provide two methods for fine-grained control over the image size input to the model:
|
||||
|
||||
1. Define min_pixels and max_pixels: Images will be resized to maintain their aspect ratio within the range of min_pixels and max_pixels.
|
||||
|
||||
2. Specify exact dimensions: Directly set `resized_height` and `resized_width`. These values will be rounded to the nearest multiple of 28.
|
||||
|
||||
```python
|
||||
# min_pixels and max_pixels
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image",
|
||||
"image": "file:///path/to/your/image.jpg",
|
||||
"resized_height": 280,
|
||||
"resized_width": 420,
|
||||
},
|
||||
{"type": "text", "text": "Describe this image."},
|
||||
],
|
||||
}
|
||||
]
|
||||
# resized_height and resized_width
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image",
|
||||
"image": "file:///path/to/your/image.jpg",
|
||||
"min_pixels": 50176,
|
||||
"max_pixels": 50176,
|
||||
},
|
||||
{"type": "text", "text": "Describe this image."},
|
||||
],
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
### Processing Long Texts
|
||||
|
||||
The current `config.json` is set for context length up to 32,768 tokens.
|
||||
To handle extensive inputs exceeding 32,768 tokens, we utilize [YaRN](https://arxiv.org/abs/2309.00071), a technique for enhancing model length extrapolation, ensuring optimal performance on lengthy texts.
|
||||
|
||||
For supported frameworks, you could add the following to `config.json` to enable YaRN:
|
||||
|
||||
```
|
||||
{
|
||||
...,
|
||||
"type": "yarn",
|
||||
"mrope_section": [
|
||||
16,
|
||||
24,
|
||||
24
|
||||
],
|
||||
"factor": 4,
|
||||
"original_max_position_embeddings": 32768
|
||||
}
|
||||
```
|
||||
|
||||
However, it should be noted that this method has a significant impact on the performance of temporal and spatial localization tasks, and is therefore not recommended for use.
|
||||
|
||||
At the same time, for long video inputs, since MRoPE itself is more economical with ids, the max_position_embeddings can be directly modified to a larger value, such as 64k.
|
||||
|
||||
|
||||
|
||||
## Citation
|
||||
|
||||
If you find our work helpful, feel free to give us a cite.
|
||||
|
||||
```
|
||||
@misc{qwen2.5-VL,
|
||||
title = {Qwen2.5-VL},
|
||||
url = {https://qwenlm.github.io/blog/qwen2.5-vl/},
|
||||
author = {Qwen Team},
|
||||
month = {January},
|
||||
year = {2025}
|
||||
}
|
||||
|
||||
@article{Qwen2VL,
|
||||
title={Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution},
|
||||
author={Wang, Peng and Bai, Shuai and Tan, Sinan and Wang, Shijie and Fan, Zhihao and Bai, Jinze and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Fan, Yang and Dang, Kai and Du, Mengfei and Ren, Xuancheng and Men, Rui and Liu, Dayiheng and Zhou, Chang and Zhou, Jingren and Lin, Junyang},
|
||||
journal={arXiv preprint arXiv:2409.12191},
|
||||
year={2024}
|
||||
}
|
||||
|
||||
@article{Qwen-VL,
|
||||
title={Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond},
|
||||
author={Bai, Jinze and Bai, Shuai and Yang, Shusheng and Wang, Shijie and Tan, Sinan and Wang, Peng and Lin, Junyang and Zhou, Chang and Zhou, Jingren},
|
||||
journal={arXiv preprint arXiv:2308.12966},
|
||||
year={2023}
|
||||
}
|
||||
```
|
||||
54
research/source-pages/qwen_license.md
Normal file
|
|
@ -0,0 +1,54 @@
|
|||
Qwen RESEARCH LICENSE AGREEMENT
|
||||
|
||||
Qwen RESEARCH LICENSE AGREEMENT Release Date: September 19, 2024
|
||||
|
||||
By clicking to agree or by using or distributing any portion or element of the Qwen Materials, you will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately.
|
||||
|
||||
1. Definitions
|
||||
a. This Qwen RESEARCH LICENSE AGREEMENT (this "Agreement") shall mean the terms and conditions for use, reproduction, distribution and modification of the Materials as defined by this Agreement.
|
||||
b. "We" (or "Us") shall mean Alibaba Cloud.
|
||||
c. "You" (or "Your") shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Materials for any purpose and in any field of use.
|
||||
d. "Third Parties" shall mean individuals or legal entities that are not under common control with us or you.
|
||||
e. "Qwen" shall mean the large language models, and software and algorithms, consisting of trained model weights, parameters (including optimizer states), machine-learning model code, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by us.
|
||||
f. "Materials" shall mean, collectively, Alibaba Cloud's proprietary Qwen and Documentation (and any portion thereof) made available under this Agreement.
|
||||
g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files.
|
||||
h. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types.
|
||||
i. "Non-Commercial" shall mean for research or evaluation purposes only.
|
||||
|
||||
2. Grant of Rights
|
||||
a. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under Alibaba Cloud's intellectual property or other rights owned by us embodied in the Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Materials FOR NON-COMMERCIAL PURPOSES ONLY.
|
||||
b. If you are commercially using the Materials, you shall request a license from us.
|
||||
|
||||
3. Redistribution
|
||||
You may distribute copies or make the Materials, or derivative works thereof, available as part of a product or service that contains any of them, with or without modifications, and in Source or Object form, provided that you meet the following conditions:
|
||||
a. You shall give any other recipients of the Materials or derivative works a copy of this Agreement;
|
||||
b. You shall cause any modified files to carry prominent notices stating that you changed the files;
|
||||
c. You shall retain in all copies of the Materials that you distribute the following attribution notices within a "Notice" text file distributed as a part of such copies: "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) Alibaba Cloud. All Rights Reserved."; and
|
||||
d. You may add your own copyright statement to your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of your modifications, or for any such derivative works as a whole, provided your use, reproduction, and distribution of the work otherwise complies with the terms and conditions of this Agreement.
|
||||
|
||||
4. Rules of use
|
||||
a. The Materials may be subject to export controls or restrictions in China, the United States or other countries or regions. You shall comply with applicable laws and regulations in your use of the Materials.
|
||||
b. If you use the Materials or any outputs or results therefrom to create, train, fine-tune, or improve an AI model that is distributed or made available, you shall prominently display “Built with Qwen” or “Improved using Qwen” in the related product documentation.
|
||||
|
||||
5. Intellectual Property
|
||||
a. We retain ownership of all intellectual property rights in and to the Materials and derivatives made by or for us. Conditioned upon compliance with the terms and conditions of this Agreement, with respect to any derivative works and modifications of the Materials that are made by you, you are and will be the owner of such derivative works and modifications.
|
||||
b. No trademark license is granted to use the trade names, trademarks, service marks, or product names of us, except as required to fulfill notice requirements under this Agreement or as required for reasonable and customary use in describing and redistributing the Materials.
|
||||
c. If you commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against us or any entity alleging that the Materials or any output therefrom, or any part of the foregoing, infringe any intellectual property or other right owned or licensable by you, then all licenses granted to you under this Agreement shall terminate as of the date such lawsuit or other proceeding is commenced or brought.
|
||||
|
||||
6. Disclaimer of Warranty and Limitation of Liability
|
||||
a. We are not obligated to support, update, provide training for, or develop any further version of the Qwen Materials or to grant any license thereto.
|
||||
b. THE MATERIALS ARE PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING WARRANTIES OF MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. WE MAKE NO WARRANTY AND ASSUME NO RESPONSIBILITY FOR THE SAFETY OR STABILITY OF THE MATERIALS AND ANY OUTPUT THEREFROM.
|
||||
c. IN NO EVENT SHALL WE BE LIABLE TO YOU FOR ANY DAMAGES, INCLUDING, BUT NOT LIMITED TO ANY DIRECT, OR INDIRECT, SPECIAL OR CONSEQUENTIAL DAMAGES ARISING FROM YOUR USE OR INABILITY TO USE THE MATERIALS OR ANY OUTPUT OF IT, NO MATTER HOW IT’S CAUSED.
|
||||
d. You will defend, indemnify and hold harmless us from and against any claim by any third party arising out of or related to your use or distribution of the Materials.
|
||||
|
||||
7. Survival and Termination.
|
||||
a. The term of this Agreement shall commence upon your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein.
|
||||
b. We may terminate this Agreement if you breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, you must delete and cease use of the Materials. Sections 6 and 8 shall survive the termination of this Agreement.
|
||||
|
||||
8. Governing Law and Jurisdiction.
|
||||
a. This Agreement and any dispute arising out of or relating to it will be governed by the laws of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement.
|
||||
b. The People's Courts in Hangzhou City shall have exclusive jurisdiction over any dispute arising out of this Agreement.
|
||||
|
||||
9. Other Terms and Conditions.
|
||||
a. Any arrangements, understandings, or agreements regarding the Material not stated herein are separate from and independent of the terms and conditions of this Agreement. You shall request a separate license from us, if you use the Materials in ways not expressly agreed to in this Agreement.
|
||||
b. We shall not be bound by any additional or different terms or conditions communicated by you unless expressly agreed.
|
||||
184
research/source-pages/smol.md
Normal file
|
|
@ -0,0 +1,184 @@
|
|||
---
|
||||
library_name: transformers
|
||||
license: apache-2.0
|
||||
datasets:
|
||||
- HuggingFaceM4/the_cauldron
|
||||
- HuggingFaceM4/Docmatix
|
||||
pipeline_tag: image-text-to-text
|
||||
language:
|
||||
- en
|
||||
base_model:
|
||||
- HuggingFaceTB/SmolLM2-360M-Instruct
|
||||
- google/siglip-base-patch16-512
|
||||
---
|
||||
|
||||
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/SmolVLM_256_banner.png" width="800" height="auto" alt="Image description">
|
||||
|
||||
# SmolVLM-500M
|
||||
|
||||
SmolVLM-500M is a tiny multimodal model, member of the SmolVLM family. It accepts arbitrary sequences of image and text inputs to produce text outputs. It's designed for efficiency. SmolVLM can answer questions about images, describe visual content, or transcribe text. Its lightweight architecture makes it suitable for on-device applications while maintaining strong performance on multimodal tasks. It can run inference on one image with 1.23GB of GPU RAM.
|
||||
|
||||
## Model Summary
|
||||
|
||||
- **Developed by:** Hugging Face 🤗
|
||||
- **Model type:** Multi-modal model (image+text)
|
||||
- **Language(s) (NLP):** English
|
||||
- **License:** Apache 2.0
|
||||
- **Architecture:** Based on [Idefics3](https://huggingface.co/HuggingFaceM4/Idefics3-8B-Llama3) (see technical summary)
|
||||
|
||||
## Resources
|
||||
|
||||
- **Demo:** [SmolVLM-256 Demo](https://huggingface.co/spaces/HuggingFaceTB/SmolVLM-256M-Demo)
|
||||
- **Blog:** [Blog post](https://huggingface.co/blog/smolvlm)
|
||||
|
||||
## Uses
|
||||
|
||||
SmolVLM can be used for inference on multimodal (image + text) tasks where the input comprises text queries along with one or more images. Text and images can be interleaved arbitrarily, enabling tasks like image captioning, visual question answering, and storytelling based on visual content. The model does not support image generation.
|
||||
|
||||
To fine-tune SmolVLM on a specific task, you can follow [the fine-tuning tutorial](https://github.com/huggingface/smollm/blob/main/vision/finetuning/Smol_VLM_FT.ipynb).
|
||||
|
||||
## Evaluation
|
||||
|
||||
|
||||
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/smoller_vlm_benchmarks.png" alt="Benchmarks" style="width:90%;" />
|
||||
|
||||
|
||||
### Technical Summary
|
||||
|
||||
SmolVLM leverages the lightweight SmolLM2 language model to provide a compact yet powerful multimodal experience. It introduces several changes compared to the larger SmolVLM 2.2B model:
|
||||
|
||||
- **Image compression:** We introduce a more radical image compression compared to Idefics3 and SmolVLM-2.2B to enable the model to infer faster and use less RAM.
|
||||
- **Visual Token Encoding:** SmolVLM-256 uses 64 visual tokens to encode image patches of size 512×512. Larger images are divided into patches, each encoded separately, enhancing efficiency without compromising performance.
|
||||
- **New special tokens:** We added new special tokens to divide the subimages. This allows for more efficient tokenization of the images.
|
||||
- **Smoller vision encoder:** We went from a 400M parameter siglip vision encoder to a much smaller 93M encoder.
|
||||
- **Larger image patches:** We are now passing patches of 512x512 to the vision encoder, instead of 384x384 like the larger SmolVLM. This allows the information to be encoded more efficiently.
|
||||
|
||||
More details about the training and architecture are available in our technical report.
|
||||
|
||||
### How to get started
|
||||
|
||||
You can use transformers to load, infer and fine-tune SmolVLM.
|
||||
|
||||
```python
|
||||
import torch
|
||||
from PIL import Image
|
||||
from transformers import AutoProcessor, AutoModelForVision2Seq
|
||||
from transformers.image_utils import load_image
|
||||
|
||||
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
|
||||
# Load images
|
||||
image = load_image("https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg")
|
||||
|
||||
# Initialize processor and model
|
||||
processor = AutoProcessor.from_pretrained("HuggingFaceTB/SmolVLM-500M-Instruct")
|
||||
model = AutoModelForVision2Seq.from_pretrained(
|
||||
"HuggingFaceTB/SmolVLM-500M-Instruct",
|
||||
torch_dtype=torch.bfloat16,
|
||||
_attn_implementation="flash_attention_2" if DEVICE == "cuda" else "eager",
|
||||
).to(DEVICE)
|
||||
|
||||
# Create input messages
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image"},
|
||||
{"type": "text", "text": "Can you describe this image?"}
|
||||
]
|
||||
},
|
||||
]
|
||||
|
||||
# Prepare inputs
|
||||
prompt = processor.apply_chat_template(messages, add_generation_prompt=True)
|
||||
inputs = processor(text=prompt, images=[image], return_tensors="pt")
|
||||
inputs = inputs.to(DEVICE)
|
||||
|
||||
# Generate outputs
|
||||
generated_ids = model.generate(**inputs, max_new_tokens=500)
|
||||
generated_texts = processor.batch_decode(
|
||||
generated_ids,
|
||||
skip_special_tokens=True,
|
||||
)
|
||||
|
||||
print(generated_texts[0])
|
||||
"""
|
||||
Assistant: The image depicts a cityscape featuring a prominent landmark, the Statue of Liberty, prominently positioned on Liberty Island. The statue is a green, humanoid figure with a crown atop its head and is situated on a small island surrounded by water. The statue is characterized by its large, detailed structure, with a statue of a woman holding a torch above her head and a tablet in her left hand. The statue is surrounded by a small, rocky island, which is partially visible in the foreground.
|
||||
In the background, the cityscape is dominated by numerous high-rise buildings, which are densely packed and vary in height. The buildings are primarily made of glass and steel, reflecting the sunlight and creating a bright, urban skyline. The skyline is filled with various architectural styles, including modern skyscrapers and older, more traditional buildings.
|
||||
The water surrounding the island is calm, with a few small boats visible, indicating that the area is likely a popular tourist destination. The water is a deep blue, suggesting that it is a large body of water, possibly a river or a large lake.
|
||||
In the foreground, there is a small strip of land with trees and grass, which adds a touch of natural beauty to the urban landscape. The trees are green, indicating that it is likely spring or summer.
|
||||
The image captures a moment of tranquility and reflection, as the statue and the cityscape come together to create a harmonious and picturesque scene. The statue's presence in the foreground draws attention to the city's grandeur, while the calm water and natural elements in the background provide a sense of peace and serenity.
|
||||
In summary, the image showcases the Statue of Liberty, a symbol of freedom and democracy, set against a backdrop of a bustling cityscape. The statue is a prominent and iconic representation of human achievement, while the cityscape is a testament to human ingenuity and progress. The image captures the beauty and complexity of urban life, with the statue serving as a symbol of hope and freedom, while the cityscape provides a glimpse into the modern world.
|
||||
"""
|
||||
```
|
||||
|
||||
|
||||
### Model optimizations
|
||||
|
||||
**Precision**: For better performance, load and run the model in half-precision (`torch.bfloat16`) if your hardware supports it.
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForVision2Seq
|
||||
import torch
|
||||
|
||||
model = AutoModelForVision2Seq.from_pretrained(
|
||||
"HuggingFaceTB/SmolVLM-Instruct",
|
||||
torch_dtype=torch.bfloat16
|
||||
).to("cuda")
|
||||
```
|
||||
|
||||
You can also load SmolVLM with 4/8-bit quantization using bitsandbytes, torchao or Quanto. Refer to [this page](https://huggingface.co/docs/transformers/en/main_classes/quantization) for other options.
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForVision2Seq, BitsAndBytesConfig
|
||||
import torch
|
||||
|
||||
quantization_config = BitsAndBytesConfig(load_in_8bit=True)
|
||||
model = AutoModelForVision2Seq.from_pretrained(
|
||||
"HuggingFaceTB/SmolVLM-Instruct",
|
||||
quantization_config=quantization_config,
|
||||
)
|
||||
```
|
||||
|
||||
**Vision Encoder Efficiency**: Adjust the image resolution by setting `size={"longest_edge": N*512}` when initializing the processor, where N is your desired value. The default `N=4` works well, which results in input images of
|
||||
size 2048×2048. Decreasing N can save GPU memory and is appropriate for lower-resolution images. This is also useful if you want to fine-tune on videos.
|
||||
|
||||
|
||||
## Misuse and Out-of-scope Use
|
||||
|
||||
SmolVLM is not intended for high-stakes scenarios or critical decision-making processes that affect an individual's well-being or livelihood. The model may produce content that appears factual but may not be accurate. Misuse includes, but is not limited to:
|
||||
|
||||
- Prohibited Uses:
|
||||
- Evaluating or scoring individuals (e.g., in employment, education, credit)
|
||||
- Critical automated decision-making
|
||||
- Generating unreliable factual content
|
||||
- Malicious Activities:
|
||||
- Spam generation
|
||||
- Disinformation campaigns
|
||||
- Harassment or abuse
|
||||
- Unauthorized surveillance
|
||||
|
||||
### License
|
||||
|
||||
SmolVLM is built upon [SigLIP](https://huggingface.co/google/siglip-base-patch16-512) as image encoder and [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct) for text decoder part.
|
||||
|
||||
We release the SmolVLM checkpoints under the Apache 2.0 license.
|
||||
|
||||
## Training Details
|
||||
|
||||
### Training Data
|
||||
|
||||
The training data comes from [The Cauldron](https://huggingface.co/datasets/HuggingFaceM4/the_cauldron) and [Docmatix](https://huggingface.co/datasets/HuggingFaceM4/Docmatix) datasets, with emphasis on document understanding (25%) and image captioning (18%), while maintaining balanced coverage across other crucial capabilities like visual reasoning, chart comprehension, and general instruction following.
|
||||
<img src="https://huggingface.co/HuggingFaceTB/SmolVLM-Instruct/resolve/main/mixture_the_cauldron.png" alt="Example Image" style="width:90%;" />
|
||||
|
||||
# Citation information
|
||||
You can cite us in the following way:
|
||||
```bibtex
|
||||
@article{marafioti2025smolvlm,
|
||||
title={SmolVLM: Redefining small and efficient multimodal models},
|
||||
author={Andrés Marafioti and Orr Zohar and Miquel Farré and Merve Noyan and Elie Bakouch and Pedro Cuenca and Cyril Zakka and Loubna Ben Allal and Anton Lozhkov and Nouamane Tazi and Vaibhav Srivastav and Joshua Lochner and Hugo Larcher and Mathieu Morlon and Lewis Tunstall and Leandro von Werra and Thomas Wolf},
|
||||
journal={arXiv preprint arXiv:2504.05299},
|
||||
year={2025}
|
||||
}
|
||||
```
|
||||
|
||||
11
research/source-pages/smolgguf.md
Normal file
|
|
@ -0,0 +1,11 @@
|
|||
---
|
||||
license: apache-2.0
|
||||
base_model: HuggingFaceTB/SmolVLM-500M-Instruct
|
||||
---
|
||||
|
||||
# SmolVLM-500M-Instruct
|
||||
|
||||
Original model: https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct
|
||||
|
||||
For more info, please refer to this PR: https://github.com/ggml-org/llama.cpp/pull/13050
|
||||
|
||||
285
research/source-pages/smolvlm2-2.2b.md
Normal file
|
|
@ -0,0 +1,285 @@
|
|||
---
|
||||
library_name: transformers
|
||||
license: apache-2.0
|
||||
datasets:
|
||||
- HuggingFaceM4/the_cauldron
|
||||
- HuggingFaceM4/Docmatix
|
||||
- lmms-lab/LLaVA-OneVision-Data
|
||||
- lmms-lab/M4-Instruct-Data
|
||||
- HuggingFaceFV/finevideo
|
||||
- MAmmoTH-VL/MAmmoTH-VL-Instruct-12M
|
||||
- lmms-lab/LLaVA-Video-178K
|
||||
- orrzohar/Video-STaR
|
||||
- Mutonix/Vript
|
||||
- TIGER-Lab/VISTA-400K
|
||||
- Enxin/MovieChat-1K_train
|
||||
- ShareGPT4Video/ShareGPT4Video
|
||||
pipeline_tag: image-text-to-text
|
||||
tags:
|
||||
- video-text-to-text
|
||||
language:
|
||||
- en
|
||||
base_model:
|
||||
- HuggingFaceTB/SmolVLM-Instruct
|
||||
---
|
||||
|
||||
|
||||
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/SmolVLM2_banner.png" width="800" height="auto" alt="Image description">
|
||||
|
||||
# SmolVLM2 2.2B
|
||||
|
||||
SmolVLM2-2.2B is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its compact size, requiring only 5.2GB of GPU RAM for video inference, it delivers robust performance on complex multimodal tasks. This efficiency makes it particularly well-suited for on-device applications where computational resources may be limited.
|
||||
## Model Summary
|
||||
|
||||
- **Developed by:** Hugging Face 🤗
|
||||
- **Model type:** Multi-modal model (image/multi-image/video/text)
|
||||
- **Language(s) (NLP):** English
|
||||
- **License:** Apache 2.0
|
||||
- **Architecture:** Based on [Idefics3](https://huggingface.co/HuggingFaceM4/Idefics3-8B-Llama3) (see technical summary)
|
||||
|
||||
## Resources
|
||||
|
||||
- **Demo:** [Video Highlight Generator](https://huggingface.co/spaces/HuggingFaceTB/SmolVLM2-HighlightGenerator)
|
||||
- **Blog:** [Blog post](https://huggingface.co/blog/smolvlm2)
|
||||
|
||||
|
||||
## Uses
|
||||
|
||||
|
||||
SmolVLM2 can be used for inference on multimodal (video / image / text) tasks where the input consists of text queries along with video or one or more images. Text and media files can be interleaved arbitrarily, enabling tasks like captioning, visual question answering, and storytelling based on visual content. The model does not support image or video generation.
|
||||
|
||||
To fine-tune SmolVLM2 on a specific task, you can follow [the fine-tuning tutorial](https://github.com/huggingface/smollm/blob/main/vision/finetuning/Smol_VLM_FT.ipynb).
|
||||
|
||||
## Evaluation
|
||||
|
||||
### Vision Evaluation
|
||||
|
||||
| Model | Mathvista | MMMU | OCRBench | MMStar | AI2D | ChartQA_Test | Science_QA | TextVQA Val | DocVQA Val |
|
||||
|-------------------|-----------|-------|----------|--------|------|--------------|------------|-------------|------------|
|
||||
| **SmolVLM2 2.2B** | 51.5 | 42 | 72.9 | 46 | 70 | 68.84 | 90 | 73.21 | 79.98 |
|
||||
| SmolVLM 2.2B | 43.9 | 38.3 | 65.5 | 41.8 | 84.5 | 71.6 | 84.5 | 72.1 | 79.7 |
|
||||
|
||||
|
||||
### Video Evaluation
|
||||
We evaluated the performance of the SmolVLM2 family on the following scientific benchmarks:
|
||||
|
||||
| Size | Video-MME | MLVU | MVBench |
|
||||
|----------|-----------------|----------|---------------|
|
||||
| 2.2B | 52.1 | 55.2 | 46.27 |
|
||||
| 500M | 42.2 | 47.3 | 39.73 |
|
||||
| 256M | 33.7 | 40.6 | 32.7 |
|
||||
|
||||
|
||||
### How to get started
|
||||
|
||||
You can use transformers to load, infer and fine-tune SmolVLM. Make sure you have num2words, flash-attn and latest transformers installed.
|
||||
You can load the model as follows.
|
||||
|
||||
```python
|
||||
from transformers import AutoProcessor, AutoModelForImageTextToText
|
||||
import torch
|
||||
|
||||
model_path = "HuggingFaceTB/SmolVLM2-2.2B-Instruct"
|
||||
processor = AutoProcessor.from_pretrained(model_path)
|
||||
model = AutoModelForImageTextToText.from_pretrained(
|
||||
model_path,
|
||||
torch_dtype=torch.bfloat16,
|
||||
_attn_implementation="flash_attention_2"
|
||||
).to("cuda")
|
||||
```
|
||||
|
||||
#### Simple Inference
|
||||
|
||||
You preprocess your inputs directly using chat templates and directly passing them
|
||||
|
||||
```python
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
|
||||
{"type": "text", "text": "Can you describe this image?"},
|
||||
]
|
||||
},
|
||||
]
|
||||
|
||||
inputs = processor.apply_chat_template(
|
||||
messages,
|
||||
add_generation_prompt=True,
|
||||
tokenize=True,
|
||||
return_dict=True,
|
||||
return_tensors="pt",
|
||||
).to(model.device, dtype=torch.bfloat16)
|
||||
|
||||
generated_ids = model.generate(**inputs, do_sample=False, max_new_tokens=64)
|
||||
generated_texts = processor.batch_decode(
|
||||
generated_ids,
|
||||
skip_special_tokens=True,
|
||||
)
|
||||
print(generated_texts[0])
|
||||
```
|
||||
|
||||
#### Video Inference
|
||||
|
||||
To use SmolVLM2 for video inference, make sure you have decord installed.
|
||||
|
||||
```python
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "video", "path": "path_to_video.mp4"},
|
||||
{"type": "text", "text": "Describe this video in detail"}
|
||||
]
|
||||
},
|
||||
]
|
||||
|
||||
inputs = processor.apply_chat_template(
|
||||
messages,
|
||||
add_generation_prompt=True,
|
||||
tokenize=True,
|
||||
return_dict=True,
|
||||
return_tensors="pt",
|
||||
).to(model.device, dtype=torch.bfloat16)
|
||||
|
||||
generated_ids = model.generate(**inputs, do_sample=False, max_new_tokens=64)
|
||||
generated_texts = processor.batch_decode(
|
||||
generated_ids,
|
||||
skip_special_tokens=True,
|
||||
)
|
||||
|
||||
print(generated_texts[0])
|
||||
```
|
||||
#### Multi-image Interleaved Inference
|
||||
|
||||
You can interleave multiple media with text using chat templates.
|
||||
|
||||
```python
|
||||
import torch
|
||||
|
||||
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": "What is the similarity between these two images?"},
|
||||
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
|
||||
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/0052a70beed5bf71b92610a43a52df6d286cd5f3/diffusers/rabbit.jpg"},
|
||||
]
|
||||
},
|
||||
]
|
||||
|
||||
inputs = processor.apply_chat_template(
|
||||
messages,
|
||||
add_generation_prompt=True,
|
||||
tokenize=True,
|
||||
return_dict=True,
|
||||
return_tensors="pt",
|
||||
).to(model.device, dtype=torch.bfloat16)
|
||||
|
||||
generated_ids = model.generate(**inputs, do_sample=False, max_new_tokens=64)
|
||||
generated_texts = processor.batch_decode(
|
||||
generated_ids,
|
||||
skip_special_tokens=True,
|
||||
)
|
||||
print(generated_texts[0])
|
||||
```
|
||||
|
||||
|
||||
### Model optimizations
|
||||
|
||||
## Misuse and Out-of-scope Use
|
||||
|
||||
SmolVLM is not intended for high-stakes scenarios or critical decision-making processes that affect an individual's well-being or livelihood. The model may produce content that appears factual but may not be accurate. Misuse includes, but is not limited to:
|
||||
|
||||
- Prohibited Uses:
|
||||
- Evaluating or scoring individuals (e.g., in employment, education, credit)
|
||||
- Critical automated decision-making
|
||||
- Generating unreliable factual content
|
||||
- Malicious Activities:
|
||||
- Spam generation
|
||||
- Disinformation campaigns
|
||||
- Harassment or abuse
|
||||
- Unauthorized surveillance
|
||||
|
||||
### License
|
||||
|
||||
SmolVLM2 is built upon [the shape-optimized SigLIP](https://huggingface.co/google/siglip-so400m-patch14-384) as image encoder and [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct) for text decoder part.
|
||||
|
||||
We release the SmolVLM2 checkpoints under the Apache 2.0 license.
|
||||
|
||||
## Citation information
|
||||
You can cite us in the following way:
|
||||
```bibtex
|
||||
@article{marafioti2025smolvlm,
|
||||
title={SmolVLM: Redefining small and efficient multimodal models},
|
||||
author={Andrés Marafioti and Orr Zohar and Miquel Farré and Merve Noyan and Elie Bakouch and Pedro Cuenca and Cyril Zakka and Loubna Ben Allal and Anton Lozhkov and Nouamane Tazi and Vaibhav Srivastav and Joshua Lochner and Hugo Larcher and Mathieu Morlon and Lewis Tunstall and Leandro von Werra and Thomas Wolf},
|
||||
journal={arXiv preprint arXiv:2504.05299},
|
||||
year={2025}
|
||||
}
|
||||
```
|
||||
|
||||
## Training Data
|
||||
SmolVLM2 used 3.3M samples for training originally from ten different datasets: [LlaVa Onevision](https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data), [M4-Instruct](https://huggingface.co/datasets/lmms-lab/M4-Instruct-Data), [Mammoth](https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M), [LlaVa Video 178K](https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K), [FineVideo](https://huggingface.co/datasets/HuggingFaceFV/finevideo), [VideoStar](https://huggingface.co/datasets/orrzohar/Video-STaR), [VRipt](https://huggingface.co/datasets/Mutonix/Vript), [Vista-400K](https://huggingface.co/datasets/TIGER-Lab/VISTA-400K), [MovieChat](https://huggingface.co/datasets/Enxin/MovieChat-1K_train) and [ShareGPT4Video](https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video).
|
||||
In the following plots we give a general overview of the samples across modalities and the source of those samples.
|
||||
<!--
|
||||
<center><img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/smolvlm2_data_split.png" width="auto" height="auto" alt="Image description">
|
||||
</center>
|
||||
|
||||
### Details
|
||||
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/smolvlm2_datadetails.png" width="auto" height="auto" alt="Image description"> -->
|
||||
|
||||
## Data Split per modality
|
||||
|
||||
| Data Type | Percentage |
|
||||
|--------------|------------|
|
||||
| Image | 34.4% |
|
||||
| Text | 20.2% |
|
||||
| Video | 33.0% |
|
||||
| Multi-image | 12.3% |
|
||||
|
||||
|
||||
## Granular dataset slices per modality
|
||||
|
||||
### Text Datasets
|
||||
| Dataset | Percentage |
|
||||
|--------------------------------------------|------------|
|
||||
| llava-onevision/magpie_pro_ft3_80b_mt | 6.8% |
|
||||
| llava-onevision/magpie_pro_ft3_80b_tt | 6.8% |
|
||||
| llava-onevision/magpie_pro_qwen2_72b_tt | 5.8% |
|
||||
| llava-onevision/mathqa | 0.9% |
|
||||
|
||||
### Multi-image Datasets
|
||||
| Dataset | Percentage |
|
||||
|--------------------------------------------|------------|
|
||||
| m4-instruct-data/m4_instruct_multiimage | 10.4% |
|
||||
| mammoth/multiimage-cap6 | 1.9% |
|
||||
|
||||
### Image Datasets
|
||||
| Dataset | Percentage |
|
||||
|--------------------------------------------|------------|
|
||||
| llava-onevision/other | 17.4% |
|
||||
| llava-onevision/vision_flan | 3.9% |
|
||||
| llava-onevision/mavis_math_metagen | 2.6% |
|
||||
| llava-onevision/mavis_math_rule_geo | 2.5% |
|
||||
| llava-onevision/sharegpt4o | 1.7% |
|
||||
| llava-onevision/sharegpt4v_coco | 1.5% |
|
||||
| llava-onevision/image_textualization | 1.3% |
|
||||
| llava-onevision/sharegpt4v_llava | 0.9% |
|
||||
| llava-onevision/mapqa | 0.9% |
|
||||
| llava-onevision/qa | 0.8% |
|
||||
| llava-onevision/textocr | 0.8% |
|
||||
|
||||
### Video Datasets
|
||||
| Dataset | Percentage |
|
||||
|--------------------------------------------|------------|
|
||||
| llava-video-178k/1-2m | 7.3% |
|
||||
| llava-video-178k/2-3m | 7.0% |
|
||||
| other-video/combined | 5.7% |
|
||||
| llava-video-178k/hound | 4.4% |
|
||||
| llava-video-178k/0-30s | 2.4% |
|
||||
| video-star/starb | 2.2% |
|
||||
| vista-400k/combined | 2.2% |
|
||||
| vript/long | 1.0% |
|
||||
| ShareGPT4Video/all | 0.8% |
|
||||
|
||||
91
scripts/ingest_training_photo.py
Executable file
|
|
@ -0,0 +1,91 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Create a metadata-free, consent-traceable Timmy training record."""
|
||||
import argparse
|
||||
import hashlib
|
||||
import hmac
|
||||
import io
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
from PIL import Image, ImageOps
|
||||
|
||||
COLORS = ("brown", "green", "yellow", "pale", "red", "black")
|
||||
QUALITIES = ("good", "fair", "poor")
|
||||
REVIEWERS = ("user-confirmed", "clinician-reviewed")
|
||||
|
||||
|
||||
def parser() -> argparse.ArgumentParser:
|
||||
p = argparse.ArgumentParser()
|
||||
p.add_argument("--input", required=True)
|
||||
p.add_argument("--output-dir", required=True)
|
||||
p.add_argument("--subject-id", required=True)
|
||||
p.add_argument("--consent-version", required=True)
|
||||
p.add_argument("--bristol-type", required=True, type=int, choices=range(1, 8))
|
||||
p.add_argument("--color", required=True, choices=COLORS)
|
||||
p.add_argument("--quality", required=True, choices=QUALITIES)
|
||||
p.add_argument("--reviewer", required=True, choices=REVIEWERS)
|
||||
return p
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parser().parse_args()
|
||||
salt = os.environ.get("TIMMY_DATASET_SALT", "")
|
||||
if len(salt) < 8:
|
||||
raise SystemExit("TIMMY_DATASET_SALT must contain at least 8 characters")
|
||||
if not args.consent_version.strip():
|
||||
raise SystemExit("consent version is required")
|
||||
|
||||
try:
|
||||
with Image.open(args.input) as source:
|
||||
image = ImageOps.exif_transpose(source).convert("RGB")
|
||||
image.thumbnail((1200, 1200))
|
||||
buffer = io.BytesIO()
|
||||
image.save(buffer, "JPEG", quality=85, optimize=True, exif=b"")
|
||||
except Exception as exc:
|
||||
raise SystemExit(f"input is not a decodable image: {exc}") from exc
|
||||
|
||||
data = buffer.getvalue()
|
||||
image_id = hashlib.sha256(data).hexdigest()
|
||||
subject_key = hmac.new(salt.encode(), args.subject_id.encode(), hashlib.sha256).hexdigest()[:16]
|
||||
bucket = int(subject_key[:2], 16)
|
||||
split = "train" if bucket < 205 else "validation" if bucket < 230 else "test"
|
||||
|
||||
output = Path(args.output_dir)
|
||||
images = output / "images"
|
||||
images.mkdir(parents=True, exist_ok=True)
|
||||
relative_path = f"images/{image_id}.jpg"
|
||||
target = output / relative_path
|
||||
if not target.exists():
|
||||
target.write_bytes(data)
|
||||
|
||||
record = {
|
||||
"schemaVersion": 1,
|
||||
"imageId": image_id,
|
||||
"relativePath": relative_path,
|
||||
"subjectKey": subject_key,
|
||||
"split": split,
|
||||
"consentVersion": args.consent_version.strip(),
|
||||
"ingestedAt": datetime.now(timezone.utc).isoformat(),
|
||||
"labels": {
|
||||
"bristolType": args.bristol_type,
|
||||
"color": args.color,
|
||||
"quality": args.quality,
|
||||
},
|
||||
"review": {"status": args.reviewer},
|
||||
"derivation": {
|
||||
"format": "jpeg",
|
||||
"maxDimension": 1200,
|
||||
"metadataStripped": True,
|
||||
},
|
||||
}
|
||||
with (output / "manifest.jsonl").open("a", encoding="utf-8") as manifest:
|
||||
manifest.write(json.dumps(record, separators=(",", ":")) + "\n")
|
||||
print(json.dumps(record))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
28
scripts/run_selfhost_smolvlm.sh
Executable file
|
|
@ -0,0 +1,28 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
LLAMA_SERVER="${LLAMA_SERVER:-/root/model-spikes/llama.cpp/build/bin/llama-server}"
|
||||
MODEL_DIR="${TIMMY_MODEL_DIR:-/root/model-spikes/models/smolvlm2-2.2b}"
|
||||
MODEL="${TIMMY_MODEL_FILE:-$MODEL_DIR/SmolVLM2-2.2B-Instruct-Q4_K_M.gguf}"
|
||||
MMPROJ="${TIMMY_MMPROJ_FILE:-$MODEL_DIR/mmproj-SmolVLM2-2.2B-Instruct-Q8_0.gguf}"
|
||||
HOST="${TIMMY_MODEL_HOST:-127.0.0.1}"
|
||||
PORT="${TIMMY_MODEL_PORT:-8080}"
|
||||
THREADS="${TIMMY_MODEL_THREADS:-4}"
|
||||
|
||||
for file in "$LLAMA_SERVER" "$MODEL" "$MMPROJ"; do
|
||||
if [[ ! -f "$file" ]]; then
|
||||
printf 'Missing required file: %s\n' "$file" >&2
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
exec "$LLAMA_SERVER" \
|
||||
--model "$MODEL" \
|
||||
--mmproj "$MMPROJ" \
|
||||
--alias SmolVLM2-2.2B-Instruct \
|
||||
--host "$HOST" \
|
||||
--port "$PORT" \
|
||||
--ctx-size 4096 \
|
||||
--parallel 1 \
|
||||
--threads "$THREADS" \
|
||||
--no-webui
|
||||
38
server.mjs
Normal file
|
|
@ -0,0 +1,38 @@
|
|||
import http from 'node:http';
|
||||
import { readFile, stat } from 'node:fs/promises';
|
||||
import { extname, join, normalize } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { analyzePhoto } from './src/vision-service.js';
|
||||
import { probeVisionProvider, resolveVisionConfig } from './src/vision-config.js';
|
||||
|
||||
const root=fileURLToPath(new URL('.',import.meta.url));
|
||||
const port=Number(process.env.PORT||4173);
|
||||
const types={'.html':'text/html; charset=utf-8','.js':'text/javascript; charset=utf-8','.css':'text/css; charset=utf-8','.json':'application/json; charset=utf-8','.webmanifest':'application/manifest+json','.svg':'image/svg+xml'};
|
||||
const visionConfig=resolveVisionConfig(process.env);
|
||||
|
||||
function sendJson(res,status,value){res.writeHead(status,{'content-type':'application/json; charset=utf-8','cache-control':'no-store','x-content-type-options':'nosniff'});res.end(JSON.stringify(value));}
|
||||
function readJson(req,maxBytes=6*1024*1024){return new Promise((resolve,reject)=>{let size=0;const chunks=[];req.on('data',chunk=>{size+=chunk.length;if(size>maxBytes){reject(new Error('Photo request is too large.'));req.destroy();return}chunks.push(chunk)});req.on('end',()=>{try{resolve(JSON.parse(Buffer.concat(chunks).toString('utf8')))}catch{reject(new Error('Invalid JSON request.'))}});req.on('error',reject)})}
|
||||
|
||||
http.createServer(async(req,res)=>{
|
||||
try{
|
||||
const url=new URL(req.url,'http://localhost');
|
||||
if(url.pathname==='/api/vision-status'&&req.method==='GET'){
|
||||
const provider=await probeVisionProvider(visionConfig);
|
||||
const privacy=visionConfig.profile==='selfhost'?'The compressed image is processed by Timmy’s self-hosted model and is not forwarded to a third-party model provider.':'A compressed copy is sent to the configured AI provider only when you explicitly request analysis.';
|
||||
return sendJson(res,200,{...visionConfig.publicStatus(),providerReady:provider.ready,modelSeen:provider.modelSeen,privacy});
|
||||
}
|
||||
if(url.pathname==='/api/analyze'&&req.method==='POST'){
|
||||
if(!visionConfig.enabled)return sendJson(res,503,{error:'AI analysis is disabled. Continue manually.'});
|
||||
const payload=await readJson(req);
|
||||
try{return sendJson(res,200,await analyzePhoto({payload,config:visionConfig}))}
|
||||
catch(error){const safe=/consent|JPEG|PNG|WebP|empty|too large/i.test(error.message);return sendJson(res,safe?400:503,{error:error.message})}
|
||||
}
|
||||
if(url.pathname.startsWith('/api/'))return sendJson(res,404,{error:'Not found'});
|
||||
if(req.method!=='GET'&&req.method!=='HEAD'){res.writeHead(405,{'allow':'GET, HEAD'});return res.end()}
|
||||
const pathname=decodeURIComponent(url.pathname);
|
||||
let path=normalize(join(root,pathname==='/'?'index.html':pathname));
|
||||
if(!path.startsWith(root))throw new Error('bad path');
|
||||
const info=await stat(path);if(info.isDirectory())path=join(path,'index.html');
|
||||
const body=await readFile(path);res.writeHead(200,{'content-type':types[extname(path)]||'application/octet-stream','cache-control':'no-store','x-content-type-options':'nosniff'});if(req.method==='HEAD')res.end();else res.end(body);
|
||||
}catch{res.writeHead(404,{'content-type':'text/plain; charset=utf-8'});res.end('Not found')}
|
||||
}).listen(port,'0.0.0.0',()=>console.log(`Timmy is listening on http://0.0.0.0:${port} · AI ${visionConfig.enabled?'ready for provider':'disabled'}`));
|
||||
5
service-worker.js
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
const CACHE='timmy-shell-v3';
|
||||
const ASSETS=['/','/index.html','/styles.css','/app.js','/src/domain.js','/src/analysis.js','/manifest.webmanifest','/assets/timmy.svg','/assets/icon-192.svg','/assets/icon-512.svg'];
|
||||
self.addEventListener('install',event=>event.waitUntil(caches.open(CACHE).then(cache=>cache.addAll(ASSETS)).then(()=>self.skipWaiting())));
|
||||
self.addEventListener('activate',event=>event.waitUntil(caches.keys().then(keys=>Promise.all(keys.filter(k=>k!==CACHE).map(k=>caches.delete(k)))).then(()=>self.clients.claim())));
|
||||
self.addEventListener('fetch',event=>{if(event.request.method!=='GET')return;event.respondWith(fetch(event.request).then(response=>{const copy=response.clone();caches.open(CACHE).then(cache=>cache.put(event.request,copy));return response}).catch(()=>caches.match(event.request).then(hit=>hit||caches.match('/index.html'))))});
|
||||
102
src/analysis.js
Normal file
|
|
@ -0,0 +1,102 @@
|
|||
const COLORS = new Set(['brown', 'green', 'yellow', 'pale', 'red', 'black']);
|
||||
const QUALITIES = new Set(['good', 'fair', 'poor']);
|
||||
const MAX_IMAGE_BYTES = 4 * 1024 * 1024;
|
||||
|
||||
function asObject(raw) {
|
||||
if (raw && typeof raw === 'object' && !Array.isArray(raw)) return raw;
|
||||
if (typeof raw !== 'string') throw new Error('Invalid AI response.');
|
||||
const cleaned = raw.trim().replace(/^```(?:json)?\s*/i, '').replace(/\s*```$/, '');
|
||||
try {
|
||||
const parsed = JSON.parse(cleaned);
|
||||
if (!parsed || typeof parsed !== 'object' || Array.isArray(parsed)) throw new Error();
|
||||
return parsed;
|
||||
} catch {
|
||||
throw new Error('Invalid AI response.');
|
||||
}
|
||||
}
|
||||
|
||||
export function parseVisionResponse(raw) {
|
||||
const value = asObject(raw);
|
||||
if (typeof value.isStool !== 'boolean') throw new Error('Invalid AI response: isStool is required.');
|
||||
const confidence = Number(value.confidence);
|
||||
if (!Number.isFinite(confidence) || confidence < 0 || confidence > 1) throw new Error('Invalid AI response: confidence is required.');
|
||||
if (!value.isStool) return { status: 'needs_user_input', isStool: false, reason: 'The image does not clearly show stool.' };
|
||||
if (confidence < 0.55) return { status: 'needs_user_input', isStool: true, reason: 'The image is too uncertain to prefill safely.' };
|
||||
const bristolType = Number(value.bristolType);
|
||||
const color = String(value.color || '').toLowerCase();
|
||||
const imageQuality = String(value.imageQuality || '').toLowerCase();
|
||||
if (!Number.isInteger(bristolType) || bristolType < 1 || bristolType > 7 || !COLORS.has(color) || !QUALITIES.has(imageQuality)) {
|
||||
throw new Error('Invalid AI response: visual fields are out of range.');
|
||||
}
|
||||
return {
|
||||
status: 'suggestion',
|
||||
isStool: true,
|
||||
bristolType,
|
||||
color,
|
||||
confidence: Math.round(confidence * 100) / 100,
|
||||
imageQuality,
|
||||
observations: String(value.observations || '').trim().slice(0, 240),
|
||||
warning: 'Visual suggestion only. Confirm it yourself; this is not a diagnosis.',
|
||||
};
|
||||
}
|
||||
|
||||
export function mergeVisualSuggestion(form, suggestion) {
|
||||
if (suggestion?.status !== 'suggestion') return { ...form };
|
||||
return { ...form, bristolType: suggestion.bristolType, color: suggestion.color };
|
||||
}
|
||||
|
||||
export function validatePhotoPayload(payload = {}) {
|
||||
if (payload.consent !== true) throw new Error('Explicit consent is required before AI analysis.');
|
||||
const match = /^data:(image\/(?:jpeg|png|webp));base64,([A-Za-z0-9+/=]+)$/.exec(String(payload.imageDataUrl || ''));
|
||||
if (!match) throw new Error('Upload a JPEG, PNG, or WebP photo.');
|
||||
const padding = (match[2].match(/=*$/) || [''])[0].length;
|
||||
const bytes = Math.floor(match[2].length * 3 / 4) - padding;
|
||||
if (bytes <= 0) throw new Error('The photo is empty.');
|
||||
if (bytes > MAX_IMAGE_BYTES) throw new Error('The photo is too large. Use an image under 4 MB.');
|
||||
return { imageDataUrl: payload.imageDataUrl, mime: match[1], bytes };
|
||||
}
|
||||
|
||||
export function buildVisionRequest({ imageDataUrl, model }) {
|
||||
return {
|
||||
model,
|
||||
temperature: 0,
|
||||
max_tokens: 500,
|
||||
messages: [{
|
||||
role: 'user',
|
||||
content: [
|
||||
{
|
||||
type: 'text',
|
||||
text: [
|
||||
'You are a conservative visual form-suggestion tool for a bowel diary, not a clinician.',
|
||||
'First decide whether the image clearly shows human stool. If it does not, set isStool=false.',
|
||||
'If it does, suggest only the closest Bristol Stool Form Scale type (1-7), visible color (brown, green, yellow, pale, red, or black), image quality, confidence, and a short neutral visual observation.',
|
||||
'Do not infer urgency, discomfort, pain, symptoms, bleeding, disease, diet safety, cause, or treatment from the image.',
|
||||
'Red or black is a visible color description only and is not a diagnosis. Be uncertain when lighting or visibility is poor.',
|
||||
'The user must confirm every suggestion. Return only the requested JSON.',
|
||||
].join(' '),
|
||||
},
|
||||
{ type: 'image_url', image_url: { url: imageDataUrl } },
|
||||
],
|
||||
}],
|
||||
response_format: {
|
||||
type: 'json_schema',
|
||||
json_schema: {
|
||||
name: 'timmy_visual_suggestion',
|
||||
strict: true,
|
||||
schema: {
|
||||
type: 'object',
|
||||
additionalProperties: false,
|
||||
required: ['isStool', 'bristolType', 'color', 'confidence', 'imageQuality', 'observations'],
|
||||
properties: {
|
||||
isStool: { type: 'boolean' },
|
||||
bristolType: { anyOf: [{ type: 'integer', minimum: 1, maximum: 7 }, { type: 'null' }] },
|
||||
color: { anyOf: [{ type: 'string', enum: [...COLORS] }, { type: 'null' }] },
|
||||
confidence: { type: 'number', minimum: 0, maximum: 1 },
|
||||
imageQuality: { type: 'string', enum: [...QUALITIES] },
|
||||
observations: { type: 'string', maxLength: 240 },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
};
|
||||
}
|
||||
75
src/domain.js
Normal file
|
|
@ -0,0 +1,75 @@
|
|||
const URGENT_KEYS = ['blood', 'blackOrDarkRed', 'severePain', 'vomiting', 'fever', 'cannotPassGas'];
|
||||
|
||||
export function bucketForBristolType(type) {
|
||||
const value = Number(type);
|
||||
if (value === 1 || value === 2) return 'constipation';
|
||||
if (value === 3 || value === 4) return 'typical';
|
||||
if (value >= 5 && value <= 7) return 'loose';
|
||||
return 'unknown';
|
||||
}
|
||||
|
||||
export function detectUrgentFlags(symptoms = {}) {
|
||||
const flags = URGENT_KEYS.filter((key) => symptoms[key] === true);
|
||||
return {
|
||||
urgent: flags.length > 0,
|
||||
flags,
|
||||
message: flags.length
|
||||
? 'These reported symptoms can need prompt medical care. Contact a clinician or urgent service now; call emergency services for heavy or nonstop bleeding, fainting, or severe worsening symptoms.'
|
||||
: 'No urgent symptom was selected. This tracker is not a diagnosis; seek care whenever you are worried or symptoms persist.',
|
||||
};
|
||||
}
|
||||
|
||||
export function buildTimmySummary(entries = []) {
|
||||
if (!entries.length) return 'No logs yet. Add one when you are ready and I’ll summarize the pattern—not diagnose it.';
|
||||
const counts = entries.reduce((acc, entry) => {
|
||||
const bucket = bucketForBristolType(entry.bristolType);
|
||||
acc[bucket] = (acc[bucket] || 0) + 1;
|
||||
return acc;
|
||||
}, {});
|
||||
const pieces = [`${entries.length} ${entries.length === 1 ? 'log' : 'logs'}`];
|
||||
if (counts.typical) pieces.push(`${counts.typical} typical`);
|
||||
if (counts.constipation) pieces.push(`${counts.constipation} on the firm side`);
|
||||
if (counts.loose) pieces.push(`${counts.loose} on the loose side`);
|
||||
return `${pieces.join(' · ')}. Patterns matter more than one entry. You choose what to eat; I only help you notice changes.`;
|
||||
}
|
||||
|
||||
export function sanitizeEntry(input = {}) {
|
||||
const symptoms = {};
|
||||
for (const key of URGENT_KEYS) symptoms[key] = input.symptoms?.[key] === true;
|
||||
const bristolType = Math.min(7, Math.max(1, Number(input.bristolType) || 4));
|
||||
return {
|
||||
id: String(input.id || globalThis.crypto?.randomUUID?.() || `${Date.now()}-${Math.random()}`),
|
||||
occurredAt: new Date(input.occurredAt || Date.now()).toISOString(),
|
||||
bristolType,
|
||||
color: ['brown', 'green', 'yellow', 'pale', 'red', 'black'].includes(input.color) ? input.color : 'brown',
|
||||
urgency: Math.min(4, Math.max(0, Number(input.urgency) || 0)),
|
||||
discomfort: Math.min(4, Math.max(0, Number(input.discomfort) || 0)),
|
||||
note: String(input.note || '').trim().slice(0, 500),
|
||||
photoDataUrl: typeof input.photoDataUrl === 'string' && input.photoDataUrl.startsWith('data:image/') ? input.photoDataUrl : '',
|
||||
symptoms,
|
||||
};
|
||||
}
|
||||
|
||||
export function photoQualityMessage({ width = 0, height = 0, brightness = 0.5 } = {}) {
|
||||
if (width < 640 || height < 480) return 'Move a little closer or use a higher-resolution photo. The image stays on this device.';
|
||||
if (brightness < 0.12) return 'Add more light before saving. Timmy only checks whether the photo is usable.';
|
||||
if (brightness > 0.95) return 'Reduce glare before saving. Timmy only checks whether the photo is usable.';
|
||||
return 'Ready for your review. Choose the matching Bristol form yourself; Timmy does not interpret the picture.';
|
||||
}
|
||||
|
||||
export function exportLedger(entries, exportedAt = new Date().toISOString()) {
|
||||
return JSON.stringify({
|
||||
product: 'Timmy the Talking Turd',
|
||||
schemaVersion: 1,
|
||||
exportedAt,
|
||||
entries: Array.isArray(entries) ? entries : [],
|
||||
}, null, 2);
|
||||
}
|
||||
|
||||
export function importLedger(text) {
|
||||
const parsed = JSON.parse(text);
|
||||
if (parsed?.schemaVersion !== 1 || !Array.isArray(parsed.entries)) throw new Error('This is not a supported Timmy export.');
|
||||
return parsed.entries.map(sanitizeEntry);
|
||||
}
|
||||
|
||||
export const urgentSymptomKeys = Object.freeze([...URGENT_KEYS]);
|
||||
58
src/vision-config.js
Normal file
|
|
@ -0,0 +1,58 @@
|
|||
const PROFILES = {
|
||||
hosted: {
|
||||
baseUrl: 'http://127.0.0.1:8645/v1',
|
||||
apiKey: 'local-proxy',
|
||||
model: 'stepfun/step-3.7-flash:free',
|
||||
processor: 'third-party',
|
||||
},
|
||||
selfhost: {
|
||||
baseUrl: 'http://127.0.0.1:8080/v1',
|
||||
apiKey: 'local-selfhost',
|
||||
model: 'SmolVLM2-2.2B-Instruct',
|
||||
processor: 'self-hosted',
|
||||
},
|
||||
};
|
||||
|
||||
export function resolveVisionConfig(env = process.env) {
|
||||
const profile = env.TIMMY_VISION_PROFILE || 'hosted';
|
||||
if (!PROFILES[profile]) throw new Error('TIMMY_VISION_PROFILE must be hosted or selfhost.');
|
||||
const defaults = PROFILES[profile];
|
||||
const config = {
|
||||
enabled: env.TIMMY_VISION_ENABLED !== '0',
|
||||
profile,
|
||||
processor: defaults.processor,
|
||||
baseUrl: env.TIMMY_VISION_BASE_URL || defaults.baseUrl,
|
||||
apiKey: env.TIMMY_VISION_API_KEY || defaults.apiKey,
|
||||
model: env.TIMMY_VISION_MODEL || defaults.model,
|
||||
requestTimeoutMs: Number(env.TIMMY_VISION_TIMEOUT_MS || (profile === 'selfhost' ? 120_000 : 60_000)),
|
||||
};
|
||||
config.publicStatus = () => ({
|
||||
enabled: config.enabled,
|
||||
profile: config.profile,
|
||||
processor: config.processor,
|
||||
model: config.enabled ? config.model : null,
|
||||
});
|
||||
return config;
|
||||
}
|
||||
|
||||
function modelsEndpoint(baseUrl) {
|
||||
const url = new URL(baseUrl);
|
||||
if (!['http:', 'https:'].includes(url.protocol)) throw new Error('Invalid AI provider URL.');
|
||||
return `${url.toString().replace(/\/$/, '')}/models`;
|
||||
}
|
||||
|
||||
export async function probeVisionProvider(config, fetchImpl = fetch) {
|
||||
if (!config.enabled) return { ready: false, modelSeen: false };
|
||||
try {
|
||||
const response = await fetchImpl(modelsEndpoint(config.baseUrl), {
|
||||
headers: { authorization: `Bearer ${config.apiKey}` },
|
||||
signal: AbortSignal.timeout(2_000),
|
||||
});
|
||||
if (!response.ok) return { ready: false, modelSeen: false };
|
||||
const body = await response.json();
|
||||
const ids = Array.isArray(body?.data) ? body.data.map(item => item?.id).filter(Boolean) : [];
|
||||
return { ready: ids.length > 0, modelSeen: ids.some(id => id === config.model || id.includes(config.model)) };
|
||||
} catch {
|
||||
return { ready: false, modelSeen: false };
|
||||
}
|
||||
}
|
||||
29
src/vision-service.js
Normal file
|
|
@ -0,0 +1,29 @@
|
|||
import { buildVisionRequest, parseVisionResponse, validatePhotoPayload } from './analysis.js';
|
||||
|
||||
function providerEndpoint(baseUrl) {
|
||||
let url;
|
||||
try { url = new URL(baseUrl); } catch { throw new Error('Invalid AI provider URL.'); }
|
||||
if (!['http:', 'https:'].includes(url.protocol)) throw new Error('Invalid AI provider URL.');
|
||||
return `${url.toString().replace(/\/$/, '')}/chat/completions`;
|
||||
}
|
||||
|
||||
export async function analyzePhoto({ payload, fetchImpl = fetch, config }) {
|
||||
const photo = validatePhotoPayload(payload);
|
||||
if (!config?.model) throw new Error('AI analysis is not configured.');
|
||||
const endpoint = providerEndpoint(config.baseUrl);
|
||||
const response = await fetchImpl(endpoint, {
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'content-type': 'application/json',
|
||||
authorization: `Bearer ${config.apiKey || 'local-proxy'}`,
|
||||
},
|
||||
body: JSON.stringify(buildVisionRequest({ imageDataUrl: photo.imageDataUrl, model: config.model })),
|
||||
signal: AbortSignal.timeout(config.requestTimeoutMs || 60_000),
|
||||
}).catch(() => { throw new Error('AI analysis is temporarily unavailable. Continue manually.'); });
|
||||
if (!response.ok) throw new Error('AI analysis is temporarily unavailable. Continue manually.');
|
||||
let data;
|
||||
try { data = await response.json(); } catch { throw new Error('Invalid response from the AI provider.'); }
|
||||
const content = data?.choices?.[0]?.message?.content;
|
||||
if (typeof content !== 'string' && (typeof content !== 'object' || content === null)) throw new Error('Invalid response from the AI provider.');
|
||||
return parseVisionResponse(content);
|
||||
}
|
||||
2
styles.css
Normal file
80
tests/analysis.test.js
Normal file
|
|
@ -0,0 +1,80 @@
|
|||
import test from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
|
||||
import {
|
||||
buildVisionRequest,
|
||||
mergeVisualSuggestion,
|
||||
parseVisionResponse,
|
||||
validatePhotoPayload,
|
||||
} from '../src/analysis.js';
|
||||
|
||||
test('validates a confident visual suggestion without inventing nonvisual fields', () => {
|
||||
const result = parseVisionResponse({
|
||||
isStool: true,
|
||||
bristolType: 4,
|
||||
color: 'brown',
|
||||
confidence: 0.82,
|
||||
imageQuality: 'good',
|
||||
observations: 'Smooth, formed appearance.',
|
||||
urgency: 4,
|
||||
discomfort: 3,
|
||||
symptoms: { blood: true },
|
||||
diagnosis: 'anything',
|
||||
});
|
||||
assert.deepEqual(result, {
|
||||
status: 'suggestion',
|
||||
isStool: true,
|
||||
bristolType: 4,
|
||||
color: 'brown',
|
||||
confidence: 0.82,
|
||||
imageQuality: 'good',
|
||||
observations: 'Smooth, formed appearance.',
|
||||
warning: 'Visual suggestion only. Confirm it yourself; this is not a diagnosis.',
|
||||
});
|
||||
assert.equal(result.urgency, undefined);
|
||||
assert.equal(result.symptoms, undefined);
|
||||
assert.equal(result.diagnosis, undefined);
|
||||
});
|
||||
|
||||
test('fails closed on a non-stool or low-confidence image', () => {
|
||||
assert.deepEqual(parseVisionResponse({ isStool: false, confidence: 0.9 }), {
|
||||
status: 'needs_user_input',
|
||||
isStool: false,
|
||||
reason: 'The image does not clearly show stool.',
|
||||
});
|
||||
assert.deepEqual(parseVisionResponse({ isStool: true, bristolType: 4, color: 'brown', confidence: 0.31 }), {
|
||||
status: 'needs_user_input',
|
||||
isStool: true,
|
||||
reason: 'The image is too uncertain to prefill safely.',
|
||||
});
|
||||
});
|
||||
|
||||
test('rejects malformed model output instead of guessing defaults', () => {
|
||||
assert.throws(() => parseVisionResponse({ isStool: true, bristolType: 9, color: 'purple', confidence: 0.8 }), /invalid/i);
|
||||
assert.throws(() => parseVisionResponse('not json'), /invalid/i);
|
||||
});
|
||||
|
||||
test('merges only visual fields and preserves user-reported context', () => {
|
||||
const form = { bristolType: 2, color: 'green', urgency: 3, discomfort: 2, note: 'user note', symptoms: { fever: true } };
|
||||
const merged = mergeVisualSuggestion(form, { status: 'suggestion', bristolType: 4, color: 'brown', confidence: 0.8 });
|
||||
assert.deepEqual(merged, { bristolType: 4, color: 'brown', urgency: 3, discomfort: 2, note: 'user note', symptoms: { fever: true } });
|
||||
});
|
||||
|
||||
test('accepts bounded JPEG/PNG/WebP data URLs and rejects oversized or unsupported input', () => {
|
||||
const payload = validatePhotoPayload({ imageDataUrl: 'data:image/jpeg;base64,' + 'YQ==', consent: true });
|
||||
assert.equal(payload.mime, 'image/jpeg');
|
||||
assert.equal(payload.bytes, 1);
|
||||
assert.throws(() => validatePhotoPayload({ imageDataUrl: 'data:image/svg+xml;base64,PHN2Zz4=', consent: true }), /JPEG, PNG, or WebP/i);
|
||||
assert.throws(() => validatePhotoPayload({ imageDataUrl: 'data:image/jpeg;base64,YQ==', consent: false }), /consent/i);
|
||||
assert.throws(() => validatePhotoPayload({ imageDataUrl: 'data:image/jpeg;base64,' + 'A'.repeat(6_000_000), consent: true }), /too large/i);
|
||||
});
|
||||
|
||||
test('builds a structured multimodal request that forbids diagnosis and nonvisual inference', () => {
|
||||
const body = buildVisionRequest({ imageDataUrl: 'data:image/jpeg;base64,YQ==', model: 'vision-model' });
|
||||
assert.equal(body.model, 'vision-model');
|
||||
assert.equal(body.response_format.type, 'json_schema');
|
||||
const prompt = body.messages[0].content.find(part => part.type === 'text').text;
|
||||
assert.match(prompt, /do not infer urgency/i);
|
||||
assert.match(prompt, /not a diagnosis/i);
|
||||
assert.equal(body.messages[0].content.find(part => part.type === 'image_url').image_url.url, 'data:image/jpeg;base64,YQ==');
|
||||
});
|
||||
88
tests/domain.test.js
Normal file
|
|
@ -0,0 +1,88 @@
|
|||
import test from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
|
||||
import {
|
||||
bucketForBristolType,
|
||||
buildTimmySummary,
|
||||
detectUrgentFlags,
|
||||
exportLedger,
|
||||
photoQualityMessage,
|
||||
sanitizeEntry,
|
||||
} from '../src/domain.js';
|
||||
|
||||
test('maps Bristol types to clinically grounded buckets', () => {
|
||||
assert.equal(bucketForBristolType(1), 'constipation');
|
||||
assert.equal(bucketForBristolType(2), 'constipation');
|
||||
assert.equal(bucketForBristolType(3), 'typical');
|
||||
assert.equal(bucketForBristolType(4), 'typical');
|
||||
assert.equal(bucketForBristolType(5), 'loose');
|
||||
assert.equal(bucketForBristolType(7), 'loose');
|
||||
assert.equal(bucketForBristolType(0), 'unknown');
|
||||
});
|
||||
|
||||
test('escalates reported blood, black stool, severe pain, vomiting, fever, or inability to pass gas', () => {
|
||||
const result = detectUrgentFlags({
|
||||
blood: true,
|
||||
blackOrDarkRed: false,
|
||||
severePain: true,
|
||||
vomiting: false,
|
||||
fever: false,
|
||||
cannotPassGas: false,
|
||||
});
|
||||
assert.equal(result.urgent, true);
|
||||
assert.deepEqual(result.flags, ['blood', 'severePain']);
|
||||
assert.match(result.message, /medical care/i);
|
||||
});
|
||||
|
||||
test('does not invent reassurance when no urgent flags are reported', () => {
|
||||
const result = detectUrgentFlags({});
|
||||
assert.equal(result.urgent, false);
|
||||
assert.deepEqual(result.flags, []);
|
||||
assert.match(result.message, /not a diagnosis/i);
|
||||
});
|
||||
|
||||
test('Timmy summary reports patterns without clearing food or diagnosing disease', () => {
|
||||
const entries = [
|
||||
{ bristolType: 3, occurredAt: '2026-08-15T08:00:00.000Z' },
|
||||
{ bristolType: 4, occurredAt: '2026-08-16T08:00:00.000Z' },
|
||||
{ bristolType: 6, occurredAt: '2026-08-17T08:00:00.000Z' },
|
||||
];
|
||||
const summary = buildTimmySummary(entries);
|
||||
assert.match(summary, /3 logs/);
|
||||
assert.match(summary, /2 typical/);
|
||||
assert.doesNotMatch(summary, /safe|diagnos|Taco Bell|clear/i);
|
||||
});
|
||||
|
||||
test('sanitizes a user entry to the MVP data contract', () => {
|
||||
const entry = sanitizeEntry({
|
||||
id: 'abc',
|
||||
occurredAt: '2026-08-17T12:00:00.000Z',
|
||||
bristolType: 4,
|
||||
color: 'brown',
|
||||
urgency: 2,
|
||||
discomfort: 1,
|
||||
note: 'After lunch',
|
||||
photoDataUrl: 'data:image/jpeg;base64,abc',
|
||||
unexpected: 'drop me',
|
||||
});
|
||||
assert.deepEqual(Object.keys(entry).sort(), [
|
||||
'bristolType', 'color', 'discomfort', 'id', 'note', 'occurredAt',
|
||||
'photoDataUrl', 'symptoms', 'urgency'
|
||||
].sort());
|
||||
assert.equal(entry.unexpected, undefined);
|
||||
});
|
||||
|
||||
test('photo quality guidance is deterministic and does not claim visual diagnosis', () => {
|
||||
assert.match(photoQualityMessage({ width: 300, height: 300, brightness: 0.5 }), /closer/i);
|
||||
assert.match(photoQualityMessage({ width: 1200, height: 900, brightness: 0.02 }), /light/i);
|
||||
assert.match(photoQualityMessage({ width: 1200, height: 900, brightness: 0.5 }), /review/i);
|
||||
assert.doesNotMatch(photoQualityMessage({ width: 1200, height: 900, brightness: 0.5 }), /type [1-7]|disease|diagnos/i);
|
||||
});
|
||||
|
||||
test('export ledger is portable JSON with version and entries', () => {
|
||||
const text = exportLedger([{ id: 'a', bristolType: 4 }], '2026-08-18T00:00:00.000Z');
|
||||
const parsed = JSON.parse(text);
|
||||
assert.equal(parsed.schemaVersion, 1);
|
||||
assert.equal(parsed.exportedAt, '2026-08-18T00:00:00.000Z');
|
||||
assert.equal(parsed.entries.length, 1);
|
||||
});
|
||||
BIN
tests/fixtures/synthetic-type4.jpg
vendored
Normal file
|
After Width: | Height: | Size: 38 KiB |
57
tests/photo-first.acceptance.mjs
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
import { chromium } from 'playwright';
|
||||
import assert from 'node:assert/strict';
|
||||
import { mkdir } from 'node:fs/promises';
|
||||
|
||||
await mkdir('artifacts', { recursive: true });
|
||||
const browser = await chromium.launch({ headless: true });
|
||||
const page = await browser.newPage({ viewport: { width: 390, height: 844 }, deviceScaleFactor: 2 });
|
||||
const errors = [];
|
||||
page.on('console', message => { if (message.type() === 'error') errors.push(message.text()); });
|
||||
page.on('pageerror', error => errors.push(error.message));
|
||||
await page.route('**/api/vision-status', route => route.fulfill({
|
||||
status: 200,
|
||||
contentType: 'application/json',
|
||||
body: JSON.stringify({ enabled: true, profile: 'selfhost', processor: 'self-hosted', model: 'SmolVLM2-2.2B-Instruct', providerReady: true, modelSeen: true }),
|
||||
}));
|
||||
await page.route('**/api/analyze', route => route.fulfill({
|
||||
status: 200,
|
||||
contentType: 'application/json',
|
||||
body: JSON.stringify({
|
||||
status: 'suggestion', isStool: true, bristolType: 4, color: 'brown', confidence: 0.83,
|
||||
imageQuality: 'good', observations: 'Smooth, formed appearance.',
|
||||
warning: 'Visual suggestion only. Confirm it yourself; this is not a diagnosis.',
|
||||
}),
|
||||
}));
|
||||
await page.goto('http://127.0.0.1:4173', { waitUntil: 'networkidle' });
|
||||
await page.evaluate(() => localStorage.clear());
|
||||
await page.reload({ waitUntil: 'networkidle' });
|
||||
|
||||
await page.locator('[data-scan]').click();
|
||||
await page.getByText(/Self-hosted model ready/i).waitFor();
|
||||
await page.screenshot({ path: 'artifacts/selfhost-photo-first-mobile.png', fullPage: false });
|
||||
assert.equal(await page.getByText('One photo. Two useful suggestions.').isVisible(), true);
|
||||
await page.locator('#ai-photo').setInputFiles('tests/fixtures/synthetic-type4.jpg');
|
||||
assert.equal(await page.locator('#analyze-photo').isDisabled(), true);
|
||||
assert.match(await page.locator('.consent-card').innerText(), /self-hosted model server/i);
|
||||
assert.doesNotMatch(await page.locator('.consent-card').innerText(), /provider’s terms/i);
|
||||
await page.locator('#ai-consent').check();
|
||||
assert.equal(await page.locator('#analyze-photo').isEnabled(), true);
|
||||
await page.locator('#analyze-photo').click();
|
||||
await page.getByText(/83% confidence/i).waitFor();
|
||||
assert.equal(await page.getByText('Type 4', { exact: true }).isVisible(), true);
|
||||
assert.equal(await page.getByText('brown', { exact: true }).isVisible(), true);
|
||||
await page.screenshot({ path: 'artifacts/photo-first-result-mobile.png', fullPage: false });
|
||||
|
||||
await page.locator('#use-suggestion').click();
|
||||
assert.equal(await page.getByText('AI PREFILLED', { exact: true }).isVisible(), true);
|
||||
assert.equal(await page.locator('[data-type="4"]').getAttribute('class').then(v => v.includes('selected')), true);
|
||||
await page.locator('#next').click();
|
||||
assert.equal(await page.locator('#color').inputValue(), 'brown');
|
||||
assert.equal(await page.locator('#urgency').inputValue(), '0');
|
||||
assert.equal(await page.locator('#discomfort').inputValue(), '0');
|
||||
assert.match(await page.locator('.fine').last().innerText(), /must come from you/i);
|
||||
await page.screenshot({ path: 'artifacts/photo-first-prefill-mobile.png', fullPage: false });
|
||||
|
||||
assert.deepEqual(errors, []);
|
||||
await browser.close();
|
||||
console.log('PASS photo → consent → AI suggestion → confirmed visual prefill → nonvisual fields remain user-reported');
|
||||
54
tests/training-ingest.test.js
Normal file
|
|
@ -0,0 +1,54 @@
|
|||
import test from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
import { mkdtemp, readFile, rm, stat } from 'node:fs/promises';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
|
||||
test('training ingest strips metadata, derives a pseudonymous split, and writes reviewable labels', async () => {
|
||||
const out = await mkdtemp(join(tmpdir(), 'timmy-ingest-'));
|
||||
try {
|
||||
const run = spawnSync('python3', [
|
||||
'scripts/ingest_training_photo.py',
|
||||
'--input', 'tests/fixtures/synthetic-type4.jpg',
|
||||
'--output-dir', out,
|
||||
'--subject-id', 'subject-123',
|
||||
'--consent-version', 'research-v1',
|
||||
'--bristol-type', '4',
|
||||
'--color', 'brown',
|
||||
'--quality', 'good',
|
||||
'--reviewer', 'user-confirmed',
|
||||
], { encoding: 'utf8', env: { ...process.env, TIMMY_DATASET_SALT: 'test-only-salt' } });
|
||||
assert.equal(run.status, 0, run.stderr);
|
||||
const record = JSON.parse(run.stdout);
|
||||
assert.match(record.imageId, /^[a-f0-9]{64}$/);
|
||||
assert.ok(['train', 'validation', 'test'].includes(record.split));
|
||||
assert.equal(record.subjectKey.length, 16);
|
||||
assert.equal(record.consentVersion, 'research-v1');
|
||||
assert.deepEqual(record.labels, { bristolType: 4, color: 'brown', quality: 'good' });
|
||||
assert.equal(record.review.status, 'user-confirmed');
|
||||
assert.equal('inputPath' in record, false);
|
||||
const derived = join(out, record.relativePath);
|
||||
assert.ok((await stat(derived)).size > 100);
|
||||
const manifest = (await readFile(join(out, 'manifest.jsonl'), 'utf8')).trim().split('\n').map(JSON.parse);
|
||||
assert.deepEqual(manifest, [record]);
|
||||
const exif = spawnSync('python3', ['-c', "from PIL import Image;import sys;print(len(Image.open(sys.argv[1]).getexif()))", derived], { encoding: 'utf8' });
|
||||
assert.equal(exif.stdout.trim(), '0');
|
||||
} finally {
|
||||
await rm(out, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('training ingest rejects records without explicit consent or human review', () => {
|
||||
const run = spawnSync('python3', [
|
||||
'scripts/ingest_training_photo.py',
|
||||
'--input', 'tests/fixtures/synthetic-type4.jpg',
|
||||
'--output-dir', '/tmp/should-not-exist-timmy',
|
||||
'--subject-id', 'subject-123',
|
||||
'--bristol-type', '4',
|
||||
'--color', 'brown',
|
||||
'--quality', 'good',
|
||||
], { encoding: 'utf8', env: { ...process.env, TIMMY_DATASET_SALT: 'test-only-salt' } });
|
||||
assert.notEqual(run.status, 0);
|
||||
assert.match(run.stderr, /consent|reviewer/i);
|
||||
});
|
||||
45
tests/ui.acceptance.mjs
Normal file
|
|
@ -0,0 +1,45 @@
|
|||
import { chromium } from 'playwright';
|
||||
import assert from 'node:assert/strict';
|
||||
import { mkdir } from 'node:fs/promises';
|
||||
|
||||
await mkdir('artifacts', { recursive: true });
|
||||
const browser = await chromium.launch({ headless: true });
|
||||
const page = await browser.newPage({ viewport: { width: 390, height: 844 }, deviceScaleFactor: 2 });
|
||||
const errors = [];
|
||||
page.on('console', message => { if (message.type() === 'error') errors.push(message.text()); });
|
||||
page.on('pageerror', error => errors.push(error.message));
|
||||
await page.goto('http://127.0.0.1:4173', { waitUntil: 'networkidle' });
|
||||
await page.evaluate(() => localStorage.clear());
|
||||
await page.reload({ waitUntil: 'networkidle' });
|
||||
|
||||
assert.match(await page.locator('h1').innerText(), /Snap first/i);
|
||||
assert.equal(await page.locator('[data-scan]').first().isVisible(), true);
|
||||
assert.equal(await page.locator('[data-log]').first().isVisible(), true);
|
||||
await page.screenshot({ path: 'artifacts/home-mobile.png', fullPage: false });
|
||||
|
||||
await page.locator('[data-log]').first().click();
|
||||
await page.locator('[data-type="6"]').click();
|
||||
await page.locator('#next').click();
|
||||
await page.locator('#color').selectOption('brown');
|
||||
await page.locator('#urgency').fill('3');
|
||||
await page.locator('#note').fill('After lunch');
|
||||
await page.locator('#next').click();
|
||||
await page.locator('[data-symptom="blood"]').check();
|
||||
assert.equal(await page.getByText('Pause and get medical help.').isVisible(), true);
|
||||
await page.screenshot({ path: 'artifacts/red-flag-mobile.png', fullPage: true });
|
||||
await page.locator('#save').click();
|
||||
assert.equal(await page.getByText('1', { exact: true }).first().isVisible(), true);
|
||||
await page.waitForTimeout(2600);
|
||||
|
||||
await page.locator('[data-view="timmy"]').last().click();
|
||||
await page.locator('[data-prompt="food"]').click();
|
||||
const chatText = await page.locator('#chat').innerText();
|
||||
assert.match(chatText, /cannot clear a restaurant/i);
|
||||
assert.doesNotMatch(chatText, /Taco Bell is safe/i);
|
||||
await page.waitForTimeout(500);
|
||||
assert.equal(await page.evaluate(() => window.scrollY), 0, 'reply should not push the page header/navigation out of frame');
|
||||
await page.screenshot({ path: 'artifacts/timmy-chat-mobile.png', fullPage: false });
|
||||
|
||||
assert.deepEqual(errors, []);
|
||||
await browser.close();
|
||||
console.log('PASS mobile home → log → red-flag → save → Timmy boundary');
|
||||
45
tests/vision-config.test.js
Normal file
|
|
@ -0,0 +1,45 @@
|
|||
import test from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
import { probeVisionProvider, resolveVisionConfig } from '../src/vision-config.js';
|
||||
|
||||
test('selfhost profile defaults to a loopback OpenAI-compatible server and no remote processor', () => {
|
||||
const config = resolveVisionConfig({ TIMMY_VISION_PROFILE: 'selfhost' });
|
||||
assert.equal(config.profile, 'selfhost');
|
||||
assert.equal(config.baseUrl, 'http://127.0.0.1:8080/v1');
|
||||
assert.equal(config.model, 'SmolVLM2-2.2B-Instruct');
|
||||
assert.equal(config.processor, 'self-hosted');
|
||||
assert.equal(config.apiKey, 'local-selfhost');
|
||||
assert.equal(config.requestTimeoutMs, 120_000);
|
||||
});
|
||||
|
||||
test('hosted profile remains explicit and never leaks its API key through public status', () => {
|
||||
const config = resolveVisionConfig({
|
||||
TIMMY_VISION_PROFILE: 'hosted',
|
||||
TIMMY_VISION_API_KEY: 'top-secret',
|
||||
});
|
||||
assert.equal(config.profile, 'hosted');
|
||||
assert.equal(config.baseUrl, 'http://127.0.0.1:8645/v1');
|
||||
assert.equal(config.apiKey, 'top-secret');
|
||||
assert.deepEqual(config.publicStatus(), {
|
||||
enabled: true,
|
||||
profile: 'hosted',
|
||||
processor: 'third-party',
|
||||
model: 'stepfun/step-3.7-flash:free',
|
||||
});
|
||||
});
|
||||
|
||||
test('provider probe checks the models endpoint and reports bounded readiness', async () => {
|
||||
let captured;
|
||||
const status = await probeVisionProvider(resolveVisionConfig({ TIMMY_VISION_PROFILE: 'selfhost' }), async (url, options) => {
|
||||
captured = { url, options };
|
||||
return { ok: true, json: async () => ({ data: [{ id: 'SmolVLM2-2.2B-Instruct' }] }) };
|
||||
});
|
||||
assert.equal(captured.url, 'http://127.0.0.1:8080/v1/models');
|
||||
assert.match(captured.options.headers.authorization, /^Bearer /);
|
||||
assert.deepEqual(status, { ready: true, modelSeen: true });
|
||||
});
|
||||
|
||||
test('provider probe fails closed without exposing upstream errors', async () => {
|
||||
const status = await probeVisionProvider(resolveVisionConfig({ TIMMY_VISION_PROFILE: 'selfhost' }), async () => { throw new Error('private upstream detail'); });
|
||||
assert.deepEqual(status, { ready: false, modelSeen: false });
|
||||
});
|
||||
48
tests/vision-service.test.js
Normal file
|
|
@ -0,0 +1,48 @@
|
|||
import test from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
import { analyzePhoto } from '../src/vision-service.js';
|
||||
|
||||
test('sends a bounded structured request to the configured provider and validates its response', async () => {
|
||||
let captured;
|
||||
const fetchImpl = async (url, options) => {
|
||||
captured = { url, options };
|
||||
return {
|
||||
ok: true,
|
||||
json: async () => ({ choices: [{ message: { content: JSON.stringify({
|
||||
isStool: true, bristolType: 4, color: 'brown', confidence: 0.81,
|
||||
imageQuality: 'good', observations: 'Smooth and formed.'
|
||||
}) } }] }),
|
||||
};
|
||||
};
|
||||
const result = await analyzePhoto({
|
||||
payload: { imageDataUrl: 'data:image/jpeg;base64,YQ==', consent: true },
|
||||
fetchImpl,
|
||||
config: { baseUrl: 'http://127.0.0.1:8645/v1', apiKey: 'secret', model: 'vision-model' },
|
||||
});
|
||||
assert.equal(result.status, 'suggestion');
|
||||
assert.equal(result.bristolType, 4);
|
||||
assert.equal(captured.url, 'http://127.0.0.1:8645/v1/chat/completions');
|
||||
assert.equal(captured.options.headers.authorization, 'Bearer secret');
|
||||
assert.doesNotMatch(JSON.stringify(result), /secret/);
|
||||
});
|
||||
|
||||
test('fails closed when the provider is unavailable or malformed', async () => {
|
||||
await assert.rejects(() => analyzePhoto({
|
||||
payload: { imageDataUrl: 'data:image/jpeg;base64,YQ==', consent: true },
|
||||
fetchImpl: async () => ({ ok: false, status: 503, text: async () => 'upstream detail' }),
|
||||
config: { baseUrl: 'http://localhost/v1', apiKey: 'x', model: 'm' },
|
||||
}), /temporarily unavailable/i);
|
||||
await assert.rejects(() => analyzePhoto({
|
||||
payload: { imageDataUrl: 'data:image/jpeg;base64,YQ==', consent: true },
|
||||
fetchImpl: async () => ({ ok: true, json: async () => ({ choices: [] }) }),
|
||||
config: { baseUrl: 'http://localhost/v1', apiKey: 'x', model: 'm' },
|
||||
}), /invalid/i);
|
||||
});
|
||||
|
||||
test('requires an http(s) provider URL and never accepts credentials from the browser payload', async () => {
|
||||
await assert.rejects(() => analyzePhoto({
|
||||
payload: { imageDataUrl: 'data:image/jpeg;base64,YQ==', consent: true, apiKey: 'browser-secret' },
|
||||
fetchImpl: async () => { throw new Error('must not call'); },
|
||||
config: { baseUrl: 'file:///tmp/provider', apiKey: 'server-secret', model: 'm' },
|
||||
}), /provider URL/i);
|
||||
});
|
||||
27
video/captions.ass
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
[Script Info]
|
||||
ScriptType: v4.00+
|
||||
PlayResX: 720
|
||||
PlayResY: 1280
|
||||
WrapStyle: 0
|
||||
ScaledBorderAndShadow: yes
|
||||
|
||||
[V4+ Styles]
|
||||
Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding
|
||||
Style: Caption,DejaVu Sans,38,&H00FFFFFF,&H000000FF,&H001D1714,&HC01D1714,-1,0,0,0,100,100,0,0,3,2,0,2,48,48,142,1
|
||||
Style: Chapter,DejaVu Sans,24,&H0043D9D0,&H000000FF,&H001D1714,&H001D1714,-1,0,0,0,100,100,2,0,3,1,0,8,34,34,60,1
|
||||
|
||||
[Events]
|
||||
Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text
|
||||
Dialogue: 0,0:00:00.40,0:00:04.60,Chapter,,0,0,0,,PRIVATE • PLAYFUL • USEFUL
|
||||
Dialogue: 0,0:00:01.00,0:00:06.00,Caption,,0,0,0,,Meet Timmy — your intelligent pooping pal.
|
||||
Dialogue: 0,0:00:06.00,0:00:10.00,Caption,,0,0,0,,A useful health habit, fast and weirdly fun.
|
||||
Dialogue: 0,0:00:10.00,0:00:14.20,Caption,,0,0,0,,Pick the closest Bristol form.
|
||||
Dialogue: 0,0:00:14.20,0:00:23.20,Caption,,0,0,0,,Add context — and an optional private photo.
|
||||
Dialogue: 0,0:00:23.20,0:00:25.20,Caption,,0,0,0,,You confirm what you saw.
|
||||
Dialogue: 0,0:00:25.20,0:00:33.80,Chapter,,0,0,0,,SAFETY FIRST
|
||||
Dialogue: 0,0:00:25.20,0:00:33.80,Caption,,0,0,0,,Red flags? Timmy stops joking.
|
||||
Dialogue: 0,0:00:33.80,0:00:38.20,Caption,,0,0,0,,Turn awkward moments into useful patterns.
|
||||
Dialogue: 0,0:00:38.20,0:00:42.20,Caption,,0,0,0,,See frequency and form over time.
|
||||
Dialogue: 0,0:00:42.20,0:00:50.20,Caption,,0,0,0,,Ask what changed — without fake diagnoses.
|
||||
Dialogue: 0,0:00:50.20,0:00:58.20,Caption,,0,0,0,,Export it. Import it. Or erase it.
|
||||
Dialogue: 0,0:00:58.20,0:01:10.80,Caption,,0,0,0,,No account. No ads. No cloud toilet gallery.
|
||||
BIN
video/final-contact.png
Normal file
|
After Width: | Height: | Size: 897 KiB |
BIN
video/hero-still.png
Normal file
|
After Width: | Height: | Size: 406 KiB |
71
video/make_music.py
Normal file
|
|
@ -0,0 +1,71 @@
|
|||
#!/usr/bin/env python3
|
||||
import wave, math
|
||||
from pathlib import Path
|
||||
import numpy as np
|
||||
|
||||
SR=48000; DURATION=76.0; BPM=100; BEAT=60/BPM
|
||||
rng=np.random.default_rng(118)
|
||||
t=np.arange(int(SR*DURATION))/SR
|
||||
left=np.zeros_like(t); right=np.zeros_like(t)
|
||||
|
||||
def add(buf,start,sound,gain=1.0):
|
||||
i=int(start*SR); j=min(len(buf),i+len(sound));
|
||||
if i<0 or i>=len(buf): return
|
||||
buf[i:j]+=sound[:j-i]*gain
|
||||
|
||||
def kick():
|
||||
d=.24; x=np.arange(int(SR*d))/SR; phase=2*np.pi*(72*x + (88-72)*(1-np.exp(-x*28))/28)
|
||||
return np.sin(phase)*np.exp(-x*18)
|
||||
def snare():
|
||||
d=.20; x=np.arange(int(SR*d))/SR; n=rng.normal(0,1,len(x)); tone=np.sin(2*np.pi*190*x)
|
||||
return (.75*n+.25*tone)*np.exp(-x*20)
|
||||
def hat(openhat=False):
|
||||
d=.17 if openhat else .055; x=np.arange(int(SR*d))/SR; n=rng.normal(0,1,len(x));
|
||||
n=np.concatenate([[0],np.diff(n)]); return n*np.exp(-x*(20 if openhat else 60))*.35
|
||||
def pluck(freq,d=.34):
|
||||
x=np.arange(int(SR*d))/SR; sig=(np.sin(2*np.pi*freq*x)+.35*np.sin(2*np.pi*2*freq*x)+.12*np.sin(2*np.pi*3*freq*x));
|
||||
return sig*np.exp(-x*8)*.25
|
||||
def bass(freq,d=.48):
|
||||
x=np.arange(int(SR*d))/SR; env=(1-np.exp(-x*30))*np.exp(-x*3.8); sig=np.sin(2*np.pi*freq*x)+.24*np.sin(2*np.pi*2*freq*x)
|
||||
return sig*env*.32
|
||||
|
||||
def pad(freq,d):
|
||||
x=np.arange(int(SR*d))/SR; env=np.minimum(1,x/.45)*np.minimum(1,(d-x)/.6); env=np.maximum(env,0)
|
||||
return (np.sin(2*np.pi*freq*x)+.22*np.sin(2*np.pi*(freq*1.003)*x))*env*.055
|
||||
|
||||
K=kick(); S=snare(); H=hat(); OH=hat(True)
|
||||
roots=[65.41,51.91,77.78,58.27] # C2 Ab1 Eb2 Bb1
|
||||
chords=[[261.63,311.13,392.00],[207.65,261.63,311.13],[155.56,196.00,233.08],[233.08,293.66,349.23]]
|
||||
# Four-bar phrases; energy opens gradually, drops under safety section, returns.
|
||||
for beat_i in range(int(DURATION/BEAT)+1):
|
||||
sec=beat_i*BEAT; bar=beat_i//4; pos=beat_i%4; section_gain=.72 if 31<=sec<43 else 1.0
|
||||
add(left,sec,K,.58*section_gain); add(right,sec,K,.58*section_gain)
|
||||
if pos in (1,3): add(left,sec,S,.18*section_gain); add(right,sec,S,.22*section_gain)
|
||||
for sub in range(2):
|
||||
hs=sec+sub*BEAT/2; pan=.86 if (beat_i*2+sub)%2 else 1.0
|
||||
add(left,hs,H,.10*pan*section_gain);add(right,hs,H,.10*(2-pan)*section_gain)
|
||||
if pos==3 and bar%2: add(left,sec+BEAT/2,OH,.07*section_gain);add(right,sec+BEAT/2,OH,.11*section_gain)
|
||||
root=roots[(bar//2)%4]
|
||||
if pos in (0,2) or (pos==3 and bar%2):
|
||||
b=bass(root*(2 if pos==2 else 1));add(left,sec,b,.75*section_gain);add(right,sec,b,.70*section_gain)
|
||||
# pads every 2 bars
|
||||
for bar in range(int(DURATION/(BEAT*4))+1):
|
||||
if bar%2: continue
|
||||
start=bar*4*BEAT; chord=chords[(bar//2)%4]
|
||||
for idx,f in enumerate(chord):
|
||||
p=pad(f,8*BEAT);add(left,start,p,.78 if idx!=1 else .58);add(right,start,p,.58 if idx!=1 else .78)
|
||||
# syncopated bright plucks outside serious segment
|
||||
melody=[392,466.16,523.25,622.25,523.25,466.16,392,349.23]
|
||||
for i in range(int(DURATION/(BEAT/2))):
|
||||
sec=i*BEAT/2
|
||||
if sec<7 or 31<=sec<43 or i%2==0: continue
|
||||
p=pluck(melody[(i//2)%len(melody)])
|
||||
add(left,sec,p,.50 if i%4==1 else .30);add(right,sec,p,.30 if i%4==1 else .50)
|
||||
# subtle stereo width and fade
|
||||
fade=np.ones_like(t);fade[:SR*2]=np.linspace(0,1,SR*2);fade[-SR*3:]=np.linspace(1,0,SR*3)
|
||||
left*=fade;right*=fade
|
||||
peak=max(np.max(np.abs(left)),np.max(np.abs(right)));left=np.tanh(left*1.25)/(max(1,peak)*1.15);right=np.tanh(right*1.25)/(max(1,peak)*1.15)
|
||||
stereo=np.stack([left,right],axis=1);pcm=(np.clip(stereo,-1,1)*32767).astype('<i2')
|
||||
out=Path(__file__).with_name('music-bed.wav')
|
||||
with wave.open(str(out),'wb') as w:w.setnchannels(2);w.setsampwidth(2);w.setframerate(SR);w.writeframes(pcm.tobytes())
|
||||
print(out)
|
||||
BIN
video/music-bed.wav
Normal file
BIN
video/narration.ogg
Normal file
BIN
video/raw-contact.png
Normal file
|
After Width: | Height: | Size: 330 KiB |
BIN
video/raw/page@2812e0e9c3ccf531013c91c700fd2e35.webm
Normal file
85
video/record_demo.mjs
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
import { chromium } from 'playwright';
|
||||
import { mkdir } from 'node:fs/promises';
|
||||
|
||||
const sleep=ms=>new Promise(r=>setTimeout(r,ms));
|
||||
await mkdir('video/raw',{recursive:true});
|
||||
const browser=await chromium.launch({headless:true});
|
||||
const context=await browser.newContext({
|
||||
viewport:{width:405,height:720},
|
||||
deviceScaleFactor:1,
|
||||
recordVideo:{dir:'video/raw',size:{width:720,height:1280}},
|
||||
colorScheme:'light'
|
||||
});
|
||||
const page=await context.newPage();
|
||||
await page.goto('http://127.0.0.1:4173',{waitUntil:'networkidle'});
|
||||
const now=Date.now();
|
||||
const seed=Array.from({length:7},(_,i)=>({
|
||||
id:`demo-${i}`,occurredAt:new Date(now-(7-i)*86400000).toISOString(),
|
||||
bristolType:[3,4,4,6,4,3,4][i],color:'brown',urgency:i===3?3:1,discomfort:i===3?2:0,
|
||||
note:['Morning','After coffee','Easy','Travel day','Back on track','Morning','Feeling regular'][i],photoDataUrl:'',
|
||||
symptoms:{blood:false,blackOrDarkRed:false,severePain:false,vomiting:false,fever:false,cannotPassGas:false}
|
||||
}));
|
||||
await page.evaluate(data=>localStorage.setItem('timmy-ledger-v1',JSON.stringify(data)),seed);
|
||||
await page.reload({waitUntil:'networkidle'});
|
||||
await page.addStyleTag({content:`
|
||||
#demo-touch{position:fixed;z-index:9999;width:54px;height:54px;border:4px solid rgba(255,255,255,.95);background:rgba(245,201,91,.5);box-shadow:0 0 0 8px rgba(21,125,120,.25);border-radius:50%;pointer-events:none;transform:translate(-50%,-50%)}
|
||||
.demo-focus{box-shadow:0 0 0 5px rgba(245,201,91,.95),0 0 0 10px rgba(21,125,120,.3)!important;transition:.2s!important}
|
||||
`});
|
||||
async function touch(selector,{click=true,after=700}={}){
|
||||
const loc=page.locator(selector).first();await loc.scrollIntoViewIfNeeded();const b=await loc.boundingBox();if(!b)throw new Error(`missing ${selector}`);
|
||||
await page.evaluate(({x,y})=>{document.querySelector('#demo-touch')?.remove();const d=document.createElement('div');d.id='demo-touch';d.style.left=x+'px';d.style.top=y+'px';document.body.append(d);d.animate([{transform:'translate(-50%,-50%) scale(.55)',opacity:.2},{transform:'translate(-50%,-50%) scale(1)',opacity:1},{transform:'translate(-50%,-50%) scale(.72)',opacity:.8}],{duration:520,easing:'ease-out'});setTimeout(()=>d.remove(),650)},{x:b.x+b.width/2,y:b.y+b.height/2});
|
||||
await sleep(300);if(click)await loc.click();await sleep(after);
|
||||
}
|
||||
async function focus(selector,ms=1500){const loc=page.locator(selector).first();await loc.scrollIntoViewIfNeeded();await loc.evaluate(e=>e.classList.add('demo-focus'));await sleep(ms);await loc.evaluate(e=>e.classList.remove('demo-focus')).catch(()=>{});}
|
||||
|
||||
// 0–8: Hero and value.
|
||||
await sleep(3000);
|
||||
await page.evaluate(()=>scrollTo({top:260,behavior:'smooth'}));await sleep(2600);
|
||||
await page.evaluate(()=>scrollTo({top:0,behavior:'smooth'}));await sleep(1400);
|
||||
// 8–17: form selection.
|
||||
await touch('[data-log]');
|
||||
await sleep(1500);
|
||||
await touch('[data-type="3"]',{after:600});
|
||||
await touch('[data-type="4"]',{after:1700});
|
||||
await touch('#next',{after:900});
|
||||
// 17–29: context and private photo lane.
|
||||
await touch('#color',{click:false,after:900});
|
||||
await page.locator('#urgency').fill('2');await focus('#urgency',900);
|
||||
await page.locator('#discomfort').fill('1');await focus('#discomfort',700);
|
||||
await page.locator('#note').fill('After lunch — feeling normal');await focus('#note',1200);
|
||||
await focus('.photo-drop',2500);
|
||||
await focus('.fine',1300);
|
||||
await touch('#next',{after:1400});
|
||||
// 29–43: safety boundary.
|
||||
await sleep(1700);
|
||||
await touch('[data-symptom="blood"]',{after:900});
|
||||
await focus('.alert',5200);
|
||||
await touch('[data-symptom="blood"]',{after:800});
|
||||
await touch('#save',{after:1700});
|
||||
// 43–51: pattern view and calendar.
|
||||
await focus('.summary-card',2500);
|
||||
await touch('[data-view="calendar"]',{after:1500});
|
||||
await focus('.calendar',3000);
|
||||
// 51–62: Timmy boundary.
|
||||
await touch('[data-view="timmy"]',{after:1600});
|
||||
await touch('[data-prompt="pattern"]',{after:1400});
|
||||
await touch('[data-prompt="food"]',{after:1100});
|
||||
await focus('#chat',3600);
|
||||
// 62–70: privacy and control.
|
||||
await touch('[data-view="privacy"]',{after:1400});
|
||||
await focus('.privacy-list',2600);
|
||||
await page.evaluate(()=>scrollTo({top:document.body.scrollHeight,behavior:'smooth'}));await sleep(2300);
|
||||
await focus('.danger-zone',1400);
|
||||
// 70–76: branded close.
|
||||
await page.evaluate(()=>{
|
||||
const o=document.createElement('div');o.id='outro';o.innerHTML=`<img src="/assets/timmy.svg"><div class="small">YOUR INTELLIGENT POOPING PAL</div><div class="big">Know your gut.<br>Keep your dignity.</div><div class="pill">TIMMY THE TALKING TURD</div>`;
|
||||
Object.assign(o.style,{position:'fixed',inset:'0',zIndex:'10000',background:'linear-gradient(145deg,#f7f3ea,#f5c95b)',display:'flex',flexDirection:'column',alignItems:'center',justifyContent:'center',textAlign:'center',color:'#28211d',fontFamily:'DM Sans,system-ui',opacity:'0'});
|
||||
o.querySelector('img').style.cssText='width:190px;height:190px;filter:drop-shadow(0 16px 18px rgba(63,37,30,.18))';
|
||||
o.querySelector('.small').style.cssText='font-size:12px;font-weight:800;letter-spacing:.13em;color:#157d78;margin-top:18px';
|
||||
o.querySelector('.big').style.cssText='font-size:39px;line-height:1.04;font-weight:800;letter-spacing:-1.5px;margin:12px 0 24px';
|
||||
o.querySelector('.pill').style.cssText='background:#28211d;color:white;padding:14px 20px;border-radius:99px;font-size:12px;font-weight:800;letter-spacing:.08em';
|
||||
document.body.append(o);o.animate([{opacity:0},{opacity:1}],{duration:700,fill:'forwards'});
|
||||
});
|
||||
await sleep(6100);
|
||||
const video=page.video();await page.close();await video.saveAs('video/timmy-app-demo-raw.webm');await context.close();await browser.close();
|
||||
console.log('video/timmy-app-demo-raw.webm');
|
||||