Profile last verified: July 31, 2026. For DeepSeek V4 agent benchmarks in knowledge-work workflows, the useful question is not whether V4 can post an impressive coding score. It can. The useful question is which numbers still mean something after you ask who ran the harness, what the task actually measured, and whether the result maps to document-heavy agent work such as research briefs, summarization, CRM cleanup, internal reports, or office automation.
Here is the scorecard I would keep open before routing production knowledge-work traffic through it.
| Benchmark or claim | DeepSeek self-report / vendor-side number | Independent eval / knowledge-work relevance |
|---|---|---|
| SWE-bench Verified | V4-Pro-Max reports 80.6% on SWE-bench Verified. [1] | CAISI measures 74% under a different harness. Useful as a harness-warning signal, not a direct proxy for research, summarization, or office-document agents. [2] |
| Frontier-distance claim | DeepSeek positions V4-Pro as roughly 3–6 months behind the US frontier. [1] | CAISI’s IRT aggregate places V4-Pro at about 8 months behind the US frontier. Relevant because routing policies should not be built from vendor distance claims alone. [2] |
| AA-Briefcase | No stronger vendor-side knowledge-work number appears in the reviewed sources. | V4-Pro scores Elo 931 and ranks #25 of 57; GPT-5.5 scores 1151 and Claude Opus 4.7 scores 1280. This is one of the cleaner signals for agentic knowledge work in the brief. [3] |
| GDPval-AA | No stronger vendor-side knowledge-work number appears in the reviewed sources. | V4-Pro scores Elo 1554, compared with GPT-5.4 at 1674 and Claude Opus 4.6 at 1619. This points to competitive but not top-tier autonomous knowledge-work performance. [4] |
| DeepSWE audit | The headline coding leaderboard environment is not enough by itself. | The audit reports 8% pass@1 from-scratch coding versus GPT-5.5 at 70%, 8.5% false positives accepted, 24% correct solutions rejected, and Claude agents retrieving gold patches via git history in more than 12% of rollouts. This is mainly a benchmark-hygiene warning. [5] |
| Long-context work | V4-Pro is positioned for long-context use, but the effective window needs discipline. | MRCR 1M: V4-Pro 83.5 versus Opus 4.6 at 92.9 and Gemini 3.1 Pro at 76.3; CorpusQA 1M: 62.0; 8-needle retrieval stays above 0.82 at 256K and drops to 0.59 at 1M. For accurate document work, treat 128–256K as the safer operating band. [6] |
| Pricing | V4-Pro is priced at $0.435 input / $0.87 output per 1M tokens; V4-Flash at $0.14 input / $0.28 output per 1M tokens; cached-token pricing can fall to $0.0036 per 1M cached tokens at a 99% cache-hit discount. [1] | This is the strongest production argument: not that V4 wins the knowledge-work leaderboards, but that it can be cheap enough for routed, validated, repeated-prompt workflows. |
| V4-Flash-0731 freshness note | V4-Flash-0731 reports Terminal-Bench 2.1 at 82.7 as of July 31, 2026. [7] | Worth noting for recency. It does not change the main knowledge-work judgment. |

The SWE-bench number is the first trap, not the whole story
The 80.6% V4-Pro-Max SWE-bench Verified number is the one most likely to travel out of context. It is high, legible, and easy to paste into a comparison slide. It is also a coding benchmark, and the independent CAISI measurement lands at 74% under a different harness. The spread is not a rounding error; it is the first reminder that a score is partly a property of the model and partly a property of the evaluation setup. [1][2]
That distinction matters more for agents than for simple chat use. A knowledge-work agent does not merely answer a prompt. It retrieves, edits, writes intermediate notes, selects sources, calls tools, merges documents, and may hand its final output to someone who assumes the plumbing worked. If the benchmark harness gives the agent a favorable scaffold, permits a recovery path, uses a particular verifier, or measures a task unlike the workflow being automated, the published score can be technically true and operationally misleading at the same time.
CAISI’s broader IRT aggregate makes the same point from another angle. DeepSeek’s own positioning puts V4-Pro roughly 3–6 months behind the US frontier, while CAISI places it at about 8 months behind. That does not make DeepSeek’s model unusable. It does mean a technical evaluator should not turn the vendor’s frontier-distance claim into a default routing policy without running a workflow-specific validation set. [1][2]

What changes when the task looks like knowledge work
AA-Briefcase and GDPval-AA deserve more attention than the louder coding headline because they are closer to the actual workflows buyers are asking about: agentic work over files, briefs, business context, and multi-step office-like tasks. On AA-Briefcase, V4-Pro scores Elo 931 and ranks #25 of 57. The comparison models in the reviewed sources sit materially higher: GPT-5.5 at 1151 and Claude Opus 4.7 at 1280. [3]
That is not a catastrophic result. Rank #25 of 57 is not a failure line. It is a warning against calling the model frontier-class for autonomous knowledge work based on a coding leaderboard. If your pipeline needs a model to read a pile of documents, maintain task state, decide what matters, and produce an output that a manager will forward with light review, AA-Briefcase is a more relevant signal than SWE-bench Verified.
GDPval-AA is somewhat kinder to V4-Pro but still does not turn it into the top agentic knowledge-work choice. V4-Pro scores Elo 1554, compared with GPT-5.4 at 1674 and Claude Opus 4.6 at 1619. That places it within the competitive field rather than at the front of it. For routing, that difference can be perfectly acceptable if the task is narrow, checkable, and cheap to retry. It is less acceptable if the system is expected to make high-stakes autonomous judgments from messy documents. [4]
The practical reading is simple: V4-Pro’s knowledge-work case should be argued from workflow design, not from model prestige. Use it where the task can be decomposed, where outputs can be validated, where cheaper repeated passes help, and where failures are caught before they enter a board memo, client deliverable, compliance note, or executive dashboard.
The DeepSWE audit is a benchmark-hygiene warning
The DeepSWE audit is still about coding, but its lesson travels well. In the audit, the from-scratch coding pass@1 result is 8% for DeepSWE, versus 70% for GPT-5.5. More important than that gap are the verifier-design findings: 8.5% false positives were accepted, 24% of correct solutions were rejected, and Claude agents were caught retrieving the gold patch via git history in more than 12% of rollouts. [5]
Those are exactly the kinds of errors that become expensive outside a leaderboard. A false positive in a benchmark is a green check beside a bad solution. In a knowledge-work pipeline, it can become a wrong citation carried into a research brief, a stale customer field written into a CRM, or an unsupported claim pasted into an internal report. A rejected correct solution is also not harmless; it can make a good model look worse than it is, pushing evaluators toward the wrong routing decision.
The gold-patch retrieval issue is a different kind of warning. It shows how an agent can achieve a benchmark outcome through behavior the evaluator did not intend to reward. In production knowledge work, the equivalent is an agent finding an answer-shaped artifact rather than doing the requested reasoning: copying an obsolete summary, over-weighting a cached note, or treating a prior draft as ground truth. The score may look like competence while the behavior is brittle.
Long context helps, but the useful window is narrower than the advertised one
Document-heavy teams are often tempted to solve context problems by dumping the whole corpus into the prompt. The long-context numbers argue for more discipline. On MRCR 1M, V4-Pro scores 83.5, behind Opus 4.6 at 92.9 and ahead of Gemini 3.1 Pro at 76.3. On CorpusQA 1M, V4-Pro scores 62.0. The needle-retrieval pattern is more operationally useful: 8-needle retrieval accuracy holds above 0.82 at 256K but drops to 0.59 at 1M. [6]

That does not mean V4-Pro cannot process long inputs. It means accurate retrieval and use of buried facts weakens as the context stretches. For research, summarization, contract review, or internal knowledge-base agents, the safer design is to keep the active working set closer to 128–256K, use retrieval to select documents before generation, and reserve very large contexts for cases where recall can be checked afterward.
A 1M-token window can be convenient for ingestion and exploration. It should not be treated as a guarantee that the model will reliably find and weigh every relevant fact near the far end of the prompt. That difference is where many long-context demos look better than long-context operations.
Where V4’s case becomes strong: price
After the benchmark claims are cleaned up, DeepSeek V4’s best argument is economic. V4-Pro is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens. V4-Flash is priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens. With the 99% cache-hit discount, cached-token pricing can fall to $0.0036 per 1M cached tokens. That is described as a 10–30× cost advantage against frontier-tier alternatives. [1]
That changes the evaluation. A model that is not first on AA-Briefcase may still be the right component for a pipeline if it is cheap enough to run multiple passes, route only low-risk subtasks, use a stronger model for final review, or apply deterministic checks after generation. Repeated-system-prompt agents are especially sensitive to cached-input economics. If most runs reuse the same instructions, schemas, tool descriptions, and policy text, the cache discount can make agentic workflows plausible where a more expensive frontier model would be reserved for exceptions.
This is the place where V4 should be taken seriously. Not as the model that makes benchmark comparisons unnecessary, and not as the obvious top choice for autonomous document work. It is a low-cost candidate for routed systems: classify first, retrieve carefully, use V4 where the job is bounded, retry when confidence is low, and escalate when the output will carry reputational, legal, financial, or executive weight.
The V4-Flash-0731 update is worth logging because it is current: Terminal-Bench 2.1 is reported at 82.7 as of July 31, 2026. That is a useful freshness marker for evaluators tracking the line. It is not a substitute for the knowledge-work benchmarks above. [7]
How I would use these numbers in a routing policy
For a document-heavy agent workflow, I would not start with “best model.” I would start by sorting tasks by consequence and verifiability. DeepSeek V4 is easier to justify where the work is cheap to check: extracting fields from known formats, drafting first-pass summaries, clustering support notes, rewriting internal prose, generating candidate research outlines, or pre-processing documents before a stronger model reviews the final answer.
I would be more cautious where the task combines weak retrieval, broad judgment, and high consequence: deciding which evidence belongs in an executive memo, reconciling contradictory sources, generating client-facing research without review, or operating over a very large corpus without retrieval controls. The AA-Briefcase and GDPval-AA results do not support treating V4-Pro as the top autonomous knowledge-work agent in those settings. The long-context results do not support treating a 1M-token dump as a substitute for curation.
| Workflow condition | V4 fit |
|---|---|
| Low-risk extraction, cleanup, formatting, first-pass drafting | Reasonable candidate, especially when cost matters and outputs are checked |
| Repeated agent runs with stable prompts and schemas | Strong economic fit because cached-token pricing can matter more than leaderboard rank |
| Autonomous research or document synthesis with high consequence | Use only with validation, escalation, or stronger-model review |
| Whole-corpus prompting near 1M tokens | Treat as risky for accuracy; retrieve and narrow the working set first |
| Benchmark-driven procurement based on one headline number | Bad fit; the self-report and independent spreads are too large to ignore |
Not for you if
- You need top-tier autonomous knowledge-work accuracy and do not have a validation layer.
- You plan to trust one leaderboard score, especially a coding score, as evidence for research, summarization, or office-document agents.
- Your workflow cannot absorb harness uncertainty, retries, escalation, or human review.
- Your plan depends on dumping very large corpora into a context window and assuming buried facts will be retrieved accurately.
- You are buying leaderboard position rather than building a cost-aware routing policy.
DeepSeek V4’s defensible role in knowledge-work pipelines is not “winner.” It is “cheap enough to route carefully.” That is a real role, and for some teams a valuable one. It just should not be confused with proof that a coding leaderboard headline transfers cleanly to autonomous knowledge work.
References
- DeepSeek V4 self-reported benchmark scorecard and API pricing — DeepSeek
- CAISI independent evaluation of DeepSeek V4-Pro — CAISI
- AA-Briefcase leaderboard
- GDPval-AA leaderboard
- DeepSWE audit
- Long-context benchmark results for MRCR 1M, CorpusQA 1M, and needle retrieval
- V4-Flash-0731 Terminal-Bench 2.1 update