An illustrated walkthrough · 36 seconds Every question is different. 01 / 06
FORGE routes evidence form and thinking depth together.
An illustrative question about the 1900 and 1924 Summer Olympics is routed to summarized evidence and a low thinking budget. A frozen host answers Paris. An easy question instead takes the Direct and NoThink route. These are explanatory examples, not measured routing predictions.
YOUR QUESTION Which city hosted both the 1900 and 1924 Summer Olympics? Which city hosted … A query, ready to route.
needs evidence?
A whole corpus. How much is enough?
① What evidence? Direct ? Query only Summary Key evidence Raw Full passages The support choice informs the thinking head.
② How much thought? NoThink CoT prompt Low budget High budget
One joint action: Summary + Low
Retrieved passages 1900 → Paris, France 1924 → Paris, France
Keep what matters.
QUERY-AWARE EXTRACT 1900 → Paris, France 1924 → Paris, France Relevant evidence, fewer tokens.
Question tokens Summary tokens Thinking: Low A tailored request
Frozen host Weights unchanged.
ANSWER TOKENS Paris [end]
QUESTION Which city hosted both the 1900 and 1924 Summer Olympics? Summary evidence + Low thinking Paris. The right support. The right amount of thought. Adapt the request around the model — keep the model frozen.
One query. One joint choice.
A DIFFERENT QUESTION How many days are in a week? No retrieved evidence needed.
FORGE Support: Direct Thinking: NoThink
Same frozen host
Answer: 7 days. Different query. Different route.
Questions differ in what they need. FORGE decides what evidence to provide and how much reasoning to use.
FORGE makes one joint choice: Direct (query only), Summary (compressed evidence), or Raw (full passages), together with a thinking setting. It sends this tailored request to a frozen host. In this illustrative example, Summary + Low yields “Paris”; an easy query can instead use Direct + NoThink.
Ⅱ Pause ↺ Replay
00:00 / 00:36
01 Question02 Joint decision03 Evidence04 Frozen host05 Answer06 Another route
Illustrative queries and choices; feature acquisition omitted. Full uses four pre-route host probes; Lite uses none. Low/High require a host with thinking-budget controls.