FROZEN AGENTS. ADAPTIVE INPUTS. PREPRINT · 2026

FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

“Choose the evidence. Set the reasoning.
Keep the model frozen.”

1 University of Alabama at Birmingham2 Tulane University3 Yale University4 Northeastern University5 Amazon AGI6 Korea University7 Carnegie Mellon University
THE IDEA, IN MOTION

A different question. A different route.

Explore the method

An illustrative walkthrough of one joint decision before the frozen host answers.

01 THE QUESTION

How much evidence?
How much reasoning?

A frozen language model can still adapt through what we put around it. But always retrieving full passages—or always thinking longer—can waste tokens and sometimes hurt accuracy.

FORGE makes the two choices together. A small router selects an evidence form and a reasoning depth for each query, balancing answer quality against input and output token cost. The host’s weights stay frozen.

269K

parameters · six-action router

42–45%

fewer cached answer tokens

8

frozen backbones · 5 benchmarks

Token savings: FORGE vs. Always-Raw on the two main hosts, with precomputed features (Table 1). Fresh Online costs are reported below.

Original FORGE overview: fixed no-retrieval and fixed retrieval workflows compared with dynamically selected evidence for a frozen language model.
FIGURE 1 From a fixed support form to a query-specific choice. View full figure ↗
02 THE METHOD

Two decisions.
One lightweight policy.

Select the support form μ and the thinking depth θ together, before generating the answer.

μ

What evidence to provide

DirectNo retrieved evidence
SummaryA query-aware extract
RawFull top-k passages
θ

How much to reason

NoThinkCoT-PromptThink-LowThink-High

6 joint actions on non-thinking hosts.
12 where thinking controls are available.

Adapt the inputs and inference procedure. Keep the host weights frozen.
FORGE router architecture: query features enter a shared MLP; a support-form head and a conditioned thinking-depth head select joint action tokens for the frozen host.
ROUTER ARCHITECTURE Shared features, two factorized heads, one joint action. View full figure ↗

Learn around the frozen host.

One cost-aware objective connects all three training stages.

  1. STAGE 0

    Enumerate

    Evaluate Direct, Summary, and Raw with NoThink on 1,500 training queries. Record answer F1 and token cost.

  2. STAGE 1

    Distill

    Fit the warm-start policy to a closed-form Boltzmann target through KL distillation.

  3. STAGE 2

    Refine

    Expand to the full joint action space, then refine with GRPO host feedback and a KL anchor.

See the three-stage training diagram
Three stages of FORGE training: offline arm enumeration, Boltzmann KL distillation, and GRPO refinement around a frozen host.
The router learns; the host stays frozen throughout.
03 THE EVIDENCE

Better answers.
More selective token use.

Five-task macro F1 on Qwen3-8B and Mistral-7B. Explore both evaluation protocols.

Cached · Qwen3-8B

Features are precomputed. Token cost counts the selected answer only; it excludes fresh feature probes.

Answer quality

Macro F1 (%) · higher is better
Always-Raw
57.1
FORGE
59.5
0100

Token cost

k tokens / query · lower is better
Always-Raw
0.430
FORGE
0.236
00.5k
+2.4 F1 pointswith 45% fewer cached answer tokens than Always-Raw.

Cached values: Table 1. Fresh Online values: Table 2, batch size 1, including every host call and input/output token. Bars share fixed zero-based scales.

CROSS-HOST TRANSFER

Train on one host.
Route for others.

A Qwen-trained Full router transfers without target-host Stage 2. It matches or exceeds Always-Raw F1 in 11 of 15 task–host cells.

Zero-shot Cached protocol. Cost excludes fresh feature probes.

Five-task macro averages · Table 8
Target hostF1
Raw → FORGE
Tokens (k)
Raw → FORGE
Llama-3.3-70B58.5 → 61.20.318 → 0.205
DeepSeek-V3.263.4 → 63.40.287 → 0.176
Qwen3.5-397B63.1 → 65.60.305 → 0.198
04 ABSTRACT

Form-optimal routing,
around a frozen model.

In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. For each query, FORGE jointly selects what evidence to provide and how much reasoning budget to allocate. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate it as a lightweight 269K-parameter factorized router. Offline arm enumeration, supervised KL distillation, and GRPO refinement train the policy around the frozen host without weight access. Across five knowledge-intensive benchmarks and eight frozen backbones, FORGE improves the quality–cost trade-off, transfers across hosts, and composes with intrinsic thinking controls where available.

Read the paper
05 CITATION

Build on FORGE.

@misc{xiao2026forge,
  title={FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents},
  author={Xi Xiao and Yunbei Zhang and Chen Liu and Lin Zhao and Jialin Chen and Tianchen Zhao and Xiang Xu and Youngeun Kim and Tianyang Wang and Min Xu},
  year={2026},
  note={Preprint},
  url={https://xixiaouab.github.io/projects/FORGE/}
}