All checks were successful
Smoke Test / smoke (pull_request) Successful in 10s
Replace Final Fantasy-inspired role codenames with generic descriptions: - Strago (Build spec author) → Build Spec: Architecture and specification - Cid (Implementation, benchmarks, deployment) → Implementation: Code, benchmarks, deployment - Locke (Research support, upstream watch) → Research: Upstream tracking, literature review - John (Quality review) → Quality: Testing and review - Frankie (Coordination) → Coordination: Project management Note: Forge URL (143.198.27.163:3000) already updated in main via PR #148 (closes #46). Refs: #67
33 lines
1.4 KiB
Markdown
33 lines
1.4 KiB
Markdown
# TurboQuant
|
|
|
|
KV cache compression for local inference on M4 Max MacBook Pro.
|
|
|
|
## What
|
|
TurboQuant (Google, ICLR 2026) is a three-stage KV cache compression method:
|
|
1. **PolarQuant** — WHT rotation + polar coordinates + Lloyd-Max codebook (~4.2x compression)
|
|
2. **QJL** — 1-bit quantized Johnson-Lindenstrauss residual correction
|
|
3. **TurboQuant** — PolarQuant + QJL = ~3.5 bits/channel, zero accuracy loss
|
|
|
|
## Why
|
|
Unlock 64K-128K context on qwen3.5:27b within 32GB unified memory.
|
|
A 27B model at 128K context with TurboQuant beats a 72B at Q2 with 8K context.
|
|
|
|
## Status
|
|
See [issues](https://forge.alexanderwhitestone.com/Timmy_Foundation/turboquant/issues) for current progress.
|
|
|
|
## Roles
|
|
- **Build Spec:** Architecture and specification
|
|
- **Implementation:** Code, benchmarks, deployment
|
|
- **Research:** Upstream tracking, literature review
|
|
- **Quality:** Testing and review
|
|
- **Coordination:** Project management
|
|
|
|
## Source Repos
|
|
- [TheTom/llama-cpp-turboquant](https://github.com/TheTom/llama-cpp-turboquant) — llama.cpp fork with Metal
|
|
- [TheTom/turboquant_plus](https://github.com/TheTom/turboquant_plus) — Reference impl, 511+ tests
|
|
- [amirzandieh/QJL](https://github.com/amirzandieh/QJL) — Author QJL code (CUDA)
|
|
- [rachittshah/mlx-turboquant](https://github.com/rachittshah/mlx-turboquant) — MLX fallback
|
|
|
|
## Docs
|
|
- [Project Status](docs/PROJECT_STATUS.md) — Full project status and build specification
|