This is super cool and very intuitively explained.
We have a wee bit of this in Braintrust: our Topics feature computes batches that reuse trace contents in KV across facets.
Now a small army of agents and I are thinking about more ways to apply this idea :)
Excited to share Quail, our new open source AI-SQL engine (a collab with Modal)! By planning queries and LLM inference together, it reaches 1B+ input tokens/min on one H100 for one query 😱🚀
AI-powered data operators create a new, interesting inference workload👇
00:00




