All checks were successful
Smoke Test / smoke (pull_request) Successful in 13s
Test matrix runner (benchmarks/run_test_matrix.py) implementing all acceptance criteria from #11: Quality Tests: - 10 practical prompts with expected-pattern matching - Perplexity proxy (WikiText-2 chunks) - Needle-in-Haystack at 8K/16K/32K contexts - Multi-turn context retention (prompt #7) Performance Tests: - tok/s at 4K/8K/16K context - TTFT proxy measurement - Peak memory (macOS/Linux) - Context ceiling binary search Outputs: - JSON: reports/test-matrix-YYYY-MM-DD.json - Markdown: reports/test-matrix-YYYY-MM-DD.md - Go/No-Go assessment with issue list Smoke test: 10/10 quality, 3/3 needle-in-haystack on qwen2.5:7b. Refs: Timmy_Foundation/turboquant#11