August 10, 2026 · Engineering record published
Measured results, published honestly
Objective: Make public technical claims traceable to real measurements — including the failures.
Completed: Long-context certification on a single 16 GB Tesla P4 — multi-needle recall 3/3 correct at 80K and 128K tokens, with the 196K+ boundary documented as a model training-context cap rather than a hardware failure. Multi-GPU engineering on the 4× P100 node — measured model comparisons (~50 vs ~31 tok/s), a 284B-parameter deployment plan with 524K-context KV planning across all four cards, and a 3× scheduler VRAM overestimate (~150 GiB claimed vs ~51 GB measured) found and documented. Thermal and lifecycle automation — idle lanes unload after 15 quiet minutes (VRAM 4,763 → 7 MiB, lowest GPU power state) and auto-wake with identity restoration and GPU-binding verification; a latent bug that had silently disabled unloads was found and fixed. Autonomous agent architecture — A2A v1.0 inter-agent transport with per-peer tokens, anti-loop caps, and signed outbound webhooks; prototype verified end-to-end in dry-run with zero API calls. Negative results published — a community inference fork measured ~8.5× slower on a live lane (4.82 vs 40.99 tok/s) and the rollout was stopped and rolled back; a Qwen3.6-35B build that emitted blank tokens was rejected.
Validation: Every figure above traces to measured results in the internal evidence store (benchmarks-20260809, kraken-parallel-priority-20260809, p4-autocool-autowake-20260809, beellama-rollout-20260809, kraken-conclave-v2-20260809). Published values are rounded and sanitized: no internal addresses, identifiers, credentials, or security-sensitive details are exposed.
Limitation: Public figures are rounded to readable values; raw measurements remain internal. Photographic evidence of the fleet is identified as an opportunity but has not been captured yet.