DEEPSEEK HARNESS TOKEN DIAGNOSTIC
Why DeepSeek Harness Uses So Many Tokens—and When MCP Tool Schemas Are the Cause
Audit the request layer before changing your stack. High usage may come from tool schemas, but tool results, history, repeated steps, reasoning, output, and cache misses are different problems with different fixes.
01 / FIND THE LAYER
Five sources can produce the same “too many tokens” symptom.
Total usage is an invoice line, not a root-cause analysis. Start with the shape of the growth.
Standing tool schemas
Every model-visible tool contributes a name, description, and input schema. DeepSeek Harness assembles registered tool schemas for each step, so a large or changing MCP catalog can dominate the request surface.
Measure: Count tools and serialize request/header.tools for each request.
Tool results and history
Arguments and model-facing result text stay in conversation history until compaction. A small catalog can still produce a large context when tools return long logs, documents, or repeated payloads.
Measure: Measure each retained message, tool argument, and tool result separately.
Repeated model steps
A turn can contain multiple model requests. Planning, a tool call, its result, and a follow-up answer each create another step, so cumulative usage is not the same metric as one request.
Measure: Count model requests per turn and compare per-request with cumulative usage.
Reasoning and output
Reasoning and final output belong to the model behavior layer. Hiding schemas does not make the model reason less or answer more briefly, and an added discovery step can increase output tokens.
Measure: Keep provider-reported output and reasoning fields separate from input.
Prompt-cache misses
A stable prefix may be reusable, while adding, removing, renaming, or changing a registered tool can alter the prefix. Cache behavior remains provider-specific and should be measured, not inferred from total input alone.
Measure: Compare cache-read and uncached input by request, then correlate changes with tool-surface changes.
02 / MEASURE IN LAYERS
A six-layer audit isolates the first source of growth.
Move from accounting to payloads, then run a controlled intervention only after the suspect layer is visible.
L0 / DEFINE
Name the number you are debugging
Record whether it is one request or the whole session, then split input, cache-read input, output, and reasoning. Do not diagnose a cumulative total as if it were one prompt.
L1 / SCHEMAS
Measure the model-facing tool surface
For every request, record tool count and serialized tool-schema JSON bytes. Compare the initial request with requests after tools/list_changed or configuration changes.
L2 / HISTORY
Measure retained messages and tool results
Attribute bytes or tokens to system prompt, user and assistant history, tool arguments, and rendered tool results. Locate the first step where growth accelerates.
L3 / STEPS
Count how often context is sent
Map each model request to the tool calls around it. A modest context sent six times can outweigh one unusually large request.
L4 / PROVIDER
Read provider usage and cache fields
Use the provider response as the source for input, cache, output, and reasoning usage. Local JSON size is a diagnostic proxy, not a token conversion.
L5 / CONTROL
Change one layer and rerun the same tasks
Hold the model, tasks, MCP server, and Harness configuration constant. Change only tool exposure; then report task completion and extra request steps beside the usage result.
request/header.tools is not a material part of the first request, a schema gateway is unlikely to fix the dominant layer.03 / CHOOSE THE RIGHT MCP PATH
Native MCP is the baseline. Progressive disclosure is a measured trade.
DeepSeek Harness is plugin-first: its documented architecture lets a plugin extend the tool registry without patching the agent loop. That makes both paths composable, but it does not make either path universally better.
OFFICIAL NATIVE MCP CLIENT
Prefer direct registration for a small, stable catalog.
- Each discovered capability is a native, server-qualified tool.
- The model has the exact schema immediately and can call it directly.
- No discovery search is required before the real call.
- Registered schema cost is present on every request while the tools remain visible.
PROGRESSIVE DISCLOSURE
Consider search-first exposure for a large, long-tail catalog.
- A stable discovery surface replaces a wall of standing schemas.
- Exact candidate schemas arrive only after a relevant search.
- Retrieval quality becomes part of task quality.
- An extra model step can raise output usage and latency.
Architecture reference: DeepSeek Harness assembles prompt sections and tool schemas at each step and exposes plugin extension points. Pinned architecture at 47f9438 ↗
04 / READ THE COMPLETE TRADE-OFF
The fixture isolates schema size. The pilot exposes the extra search.
FIXED 1,000-TOOL COMPONENT FIXTURE
647,962 B → 1,114 B
Serialized registered tool-schema JSON bytes only. This is not a provider-token, latency, or task-quality measurement.
Inspect the rc.9 benchmark ↗THREE-TASK DEEPSEEK HARNESS PILOT
3 / 3 ↔ 3 / 3
Both arms completed all three fixed tasks. Direct MCP made the selected tool call; Lens first called mcp_search, then mcp_call.
| Metric | Official direct client | MCP Lens |
|---|---|---|
| Completed tasks | 3 / 3 | 3 / 3 |
| MCP path | Direct selected-tool call | mcp_search → mcp_call |
| Output tokens | 491 | 794 |
Interpretation: those three cases show task-completion parity and a smaller standing schema surface, but Lens paid for discovery with another call and higher output usage. They do not establish a general quality or latency result.
05 / APPLY THE NARROW FIX
What MCP Lens changes—and what it leaves untouched.
LENS DOES
- Keep the model-facing MCP surface at
mcp_searchandmcp_call. - Return a bounded candidate set with exact schemas after search.
- Call one explicit server and tool after another policy check.
- Start with
allowTools: []so remote capability exposure is opt-in.
LENS DOES NOT
- Shrink tool-result text or conversation history already retained by Harness.
- Reduce model reasoning or final-answer length.
- Fix provider cache misses unrelated to the tool surface.
- Guarantee lower tokens, cost, latency, or better task quality.
- Sandbox remote MCP processes or replace endpoint security.
NEXT STEP
Measure first. If schemas dominate, test a reversible profile.
The homepage has a copy-paste DeepSeek Harness profile, default-deny allowlist, validation command, and first test prompt.
FAQ / DIRECT ANSWERS
DeepSeek Harness token usage, without the shortcuts.
01Why is DeepSeek Harness using so many tokens?
High usage can come from standing tool schemas, retained tool results and history, repeated model steps, reasoning and output, or prompt-cache misses. Split those layers before changing the MCP setup.
02How can I tell whether MCP tool schemas are the cause?
Measure the tool count and serialized request/header.tools JSON on each request. If that surface is large before any tool result exists, and a controlled run with fewer visible tools reduces the same layer, schema bloat is a plausible cause.
03Does the official DeepSeek Harness MCP client put schemas in each request?
The pinned official client documentation says discovered tools are registered as native tools and that their data-dependent schema cost is paid on every request while they remain registered.
04When should I keep the official native MCP client?
Keep it for a small, stable catalog where direct calls and the simplest execution path matter more than reducing the standing tool surface. It avoids a separate discovery call.
05When does progressive disclosure make sense?
Consider it for dozens or hundreds of tools, multiple servers, long-tail capabilities, or frequent catalog changes—after measuring that the standing schema layer is material. Account for retrieval quality and the added search step.
06Does MCP Lens guarantee lower total token usage?
No. The fixed fixture measures schema JSON bytes, not tokens. In the three-task pilot both arms completed 3/3, while Lens added a search step and used 794 output tokens versus 491 for the direct client.