turboquant

Timmy_Foundation/turboquant

Fork 0

Files

History

PRIMA 02c0cc2b23

Smoke Test / smoke (pull_request) Successful in 15s

Details

test: tool call regression suite for compressed models (closes #96 )

tests/tool_call_regression.py:
- 10 test cases covering 5 hermes tools: read_file, web_search, terminal,
  execute_code, delegate_task
- Schema validation (OpenAI-compatible tool call format)
- Argument validation (correct tool + expected args)
- Parallel tool calling test (multiple tools in one response)
- Dry-run mode for CI (schema validation without server)
- Full server mode with latency tracking
- Markdown report generation with results matrix
- JSON results output for programmatic consumption
- 95% accuracy threshold gate (exit code 1 on failure)

benchmarks/tool-call-regression.md:
- Results template with model/preset matrix
- Tool coverage tracking table

.gitea/workflows/smoke.yml:
- Added dry-run tool call schema validation step

2026-04-15 21:58:34 -04:00

perplexity_results.json

feat: wikitext-2 corpus + perplexity benchmark script (closes #21 )

2026-04-12 00:39:14 -04:00

prompts.json

feat: add standardized benchmarking prompts

2026-03-30 21:14:48 +00:00

run_benchmarks.py

feat: multi-backend benchmark suite with TTFT + memory tracking (#37 )

2026-04-13 14:05:17 +00:00

run_long_session.py

burn: add long-session quality test (Issue #12 ) (#39 )