01 · Context distillation
Instructions, conversation history, retrieved documents and tool scaffolding repeat on nearly every interaction. TokenTrim collapses that redundant footprint, prompt and KV compression working together, so the model sees the signal, not the bulk. Pre-inference, in-line, before a single token reaches the model.
