DeepSeek Harness (dsh): Ultimate Guide to AI Agent Harnesses
A frontier language model by itself cannot open your repository, decide which files matter, run a test, notice that the test failed, repair the code, ask before touching a sensitive path, remember what happened, and resume the work tomorrow.
The harness is the software around the model that makes those things possible.
That distinction has become important enough that comparing “Model A versus Model B” is often incomplete. Two products can use the same model and still behave very differently because their tools, prompts, context management, agent loop, sandbox, retry logic and verification strategy are different.
DeepSeek Harness, usually shortened to dsh, makes that surrounding layer unusually explicit. DeepSeek describes the product with a simple equation:
Agent = Model + Harness
Its more distinctive idea is “everything is a plugin.” Models, tools, skills, sessions, sandboxes, storage, agent loops, scheduling and even the UI are designed as replaceable parts.
This guide explains what an AI agent harness actually is, why it can materially change the performance of the same model, how DeepSeek Harness works, how to install and use it safely, where it fits against Claude Code, Codex CLI, Gemini CLI, OpenCode and OpenHands, and when you should use a harness rather than an agent framework.
Current status, August 21, 2026: DeepSeek Harness is open source under MIT but is explicitly a developer preview. DeepSeek warns that compatibility-breaking changes will occur. It is compelling infrastructure for experimentation and custom agent systems, but it should not yet be treated as a stable production platform.
Current project snapshot
| Item | Status on August 21, 2026 |
|---|---|
| Project | DeepSeek Harness (dsh) |
| License | MIT |
| Maturity | Developer preview; breaking changes explicitly expected |
| Fast start | npx @deepseek-ai/dsh web |
| Default UI | http://127.0.0.1:3080 |
| Source workspace version | 0.1.0-rc.8 on current repository master |
| npm channel at review time | latest was 0.1.0-rc.7; rc.8 was on the newer channel |
| Main surfaces | Web UI, headless execution, Python SDK |
| Default security posture | Loopback Web UI; workspace-write + approval-oriented permissions |
| Production recommendation | Evaluate and pin versions; do not assume API/plugin compatibility yet |
The exact release tag is less important than the maturity label. dsh is moving fast enough that repository master, npm channels and documentation can briefly differ. Pin the version or commit whenever repeatability matters.
The shortest possible explanation
Think of an AI model as an engine and driver brain. It can reason and produce instructions, but it does not automatically have wheels, brakes, mirrors, navigation, a garage or traffic rules.
The harness supplies those surrounding systems.
User goal
|
v
+---------------------------+
| Agent Harness |
| |
| instructions / context |
| agent loop |
| tools and tool schemas |
| files / shell / web |
| memory and sessions |
| approvals / sandbox |
| retries / verification |
| logs / traces / UI |
+-------------+-------------+
|
v
Language model
|
v
tool calls / answer
|
+----> back through harness
A normal chatbot mostly follows:
prompt -> model -> answer
An agent follows something closer to:
goal -> model -> tool -> observation -> model -> tool -> verification -> ... -> answer
The harness owns most of that second loop.
What exactly is an AI agent harness?
The term is still used inconsistently. It can refer to a complete coding-agent product such as Claude Code or Codex CLI, the runtime underneath an agent, or in other contexts an evaluation harness that executes benchmark tasks.
For practical engineering, an agent harness is the runtime layer that wraps a model and repeatedly connects it to a real environment until a task is complete or the run must stop.
A useful harness normally handles at least several of these responsibilities:
| Harness responsibility | Why it exists |
|---|---|
| System instructions | Tells the model its role, constraints and operating rules |
| Context assembly | Decides which conversation, files, tool results and state enter the next model request |
| Agent loop | Decides when to call the model again and when a turn is finished |
| Tool registry | Defines actions such as read, edit, shell, search, browser or database access |
| Tool execution | Actually performs a requested action and returns structured results |
| Sessions / memory | Preserves useful state across steps or future runs |
| Sandboxing | Restricts what agent-created processes can access |
| Approval policy | Decides which actions require a human decision |
| Verification | Runs tests, linters, checks or other evidence-producing operations |
| Recovery | Handles failed tools, retries, timeouts and partial progress |
| Subagents / delegation | Allows work to be split among specialized agents |
| Observability | Records prompts, actions, results, timings and failures |
| UI / API | Gives humans or other software a way to drive the agent |
| Scheduling | Runs work later or repeatedly without an interactive UI |
This definition separates the harness from several related concepts.
Harness vs model
A model predicts and reasons over tokens and may emit tool calls. A harness decides what tools exist, executes them, returns observations, manages context and controls the run.
Changing the model can change intelligence. Changing the harness can change what that intelligence is able to see, do, remember, verify and recover from.
Harness vs MCP
Model Context Protocol (MCP) is a protocol for connecting AI applications to tools and data sources. MCP is comparable to a standardized connector.
A harness is much larger. It may use MCP, but it also owns the loop, permissions, context, sessions, tool presentation, sandbox and other runtime behavior.
MCP server != agent harness
Harness vs skill
A skill is usually reusable instructions, scripts or resources for a particular task. Skills teach or equip an agent for a domain.
The harness decides how skills are discovered, loaded, authorized and combined with the rest of the runtime.
Harness vs agent framework or SDK
Frameworks such as LangGraph, AutoGen, CrewAI and agent SDKs are primarily building blocks for developers constructing agents or workflows.
A product-style harness such as dsh, Claude Code or Codex CLI is closer to a ready runtime: it already has an agent loop, working tools, session behavior, approvals and user-facing execution surfaces.
The boundary can blur. OpenHands, for example, exposes both applications and a composable SDK. DeepSeek Harness itself is deliberately designed to be infrastructure that developers can reshape.
Harness vs evaluation harness
An evaluation harness runs controlled tasks, records results and calculates benchmark metrics. It does not necessarily turn a model into a useful day-to-day agent.
DeepSeek's Minimal mode is useful precisely because an agent harness can also be configured as a minimal environment for cleaner model evaluation.
Why harnesses suddenly matter so much
Earlier LLM applications were dominated by the quality of the prompt and the model. Agentic workloads add another major variable: interaction with the environment.
A coding agent may need to:
- inspect a repository;
- locate the relevant implementation;
- form a hypothesis;
- edit multiple files;
- run tests;
- interpret an error;
- search for an API contract;
- revise the patch;
- run focused and broad validation;
- stop without undoing successful work.
A weak harness can waste an excellent model by feeding huge irrelevant outputs, presenting ambiguous tools, failing to preserve state, allowing loops to spin, hiding sandbox failures or stopping before verification.
A strong harness can make the same model appear dramatically more capable because it converts reasoning into useful feedback and action.
That does not mean a harness creates intelligence from nothing. It cannot give a weak model knowledge or reasoning ability it fundamentally lacks. It controls how efficiently the model's existing capability reaches the environment.
Research: how much can the harness change model performance?
This is no longer only an intuition from coding-agent users.
A June 2026 study, “The Scaffold Effect in Coding Agents,” ran the same Qwen 3.6 Plus and MiniMax M2.5 models through Goose, OpenCode and OpenHands SDK on a 50-task Terminal-Bench Pro subset. Harness choice produced as much as a 40× difference in tokens per solved task. Within-model pass-rate changes were smaller in that experiment, 0–8 percentage points, but failure patterns were strongly harness-specific.
That finding is important for two reasons:
- efficiency can change enormously even when raw success rate moves only modestly;
- a leaderboard row named only after the model can hide the runtime that shaped the result.
First-party engineering results point in the same direction, though they should be read as product-team evaluations rather than neutral benchmarks. LangChain reported taking the same gpt-5.2-codex model from 52.8% to 66.5% on Terminal-Bench 2.0 by changing the harness layer, including prompts and middleware. In a later model-profile experiment, it reported moving GPT-5.3 Codex from 33% to 53% and Claude Opus 4.7 from 43% to 53% on a difficult tau2-bench subset by tuning harness behavior per model.
The two results are complementary: the Scaffold Effect study shows that harness choice can radically change efficiency and failure shape, while model-specific harness tuning shows that prompt/tool/middleware fit can also move task success materially.
Another 2026 line of research on effective feedback compute argues that success depends less on simply spending more tokens or tool calls and more on whether the harness turns those interactions into valid, informative, non-redundant feedback that survives into later decisions.
The practical lesson is:
Evaluate the model-harness pair, not the model in isolation.
For a real deployment, track at least:
- task success rate;
- tokens and cost per successful task;
- wall-clock latency;
- number of tool calls;
- recovery after tool failure;
- unnecessary actions;
- verification quality;
- approval burden; and
- security boundary violations.
What DeepSeek Harness is
DeepSeek Harness is an open-source agent runtime from DeepSeek AI. Its command-line launcher is dsh.
The project is built on Cordis, a plugin system in which services, typed events and reversible effects are mounted into a shared context. DeepSeek's architecture documentation states that every part of the product is a plugin, including:
- model adapters;
- tools;
- skills;
- session logging;
- storage;
- sandboxing;
- approval policy;
- agent presets;
- agent loops;
- web access;
- subagents;
- jobs and scheduling; and
- UI integrations.
The unusual claim is not merely “we support plugins.” The agent loop itself is replaceable.
That makes dsh closer to a programmable agent runtime or small agent operating environment than a fixed terminal assistant with an extension API bolted on.
The core design: “everything is a plugin”
Many applications have a hard-coded core and allow extensions around the edges.
DeepSeek takes a different approach. A running dsh instance is composed from plugins and configuration. The Cordis kernel manages mounting, unmounting and dependencies.
The official architecture uses capability seams. A capability is split conceptually into:
- Service definition – the interface;
- Service provider – an implementation; and
- Consumer – something that uses the capability.
For example, a file-editing tool does not have to contain all the logic for local files, a remote sandbox and an E2B-style backend. It can consume a filesystem capability while different providers implement the environment.
This is a powerful architectural property because it reduces coupling.
model-facing file tool
|
v
filesystem capability
/ | \
local sandboxed remote
The same principle applies to model providers, shells, web access, subagents and other runtime components.
Profiles, bundles and presets: three terms worth understanding
DeepSeek Harness has several composition layers. They sound similar but solve different problems.
Profile
A profile describes how the whole dsh process boots.
DeepSeek currently ships profile templates including web and headless.
The base layer supplies capabilities such as model adapters, persistence, tools, sandboxing, approvals, credentials and telemetry. A web profile adds the HTTP/browser application. A headless profile adds a one-shot runner without opening a web server.
You can inspect the effective composition with:
dsh --profile web --dump-config
Bundle
A bundle packages Cordis configuration rows and the code they mount. Profiles stack bundles in order, then user patches can override the composition.
Bundles are distribution units for runtime configuration.
Agent preset
An agent preset changes what an individual session's agent contains: its persona, tools, prompt sections and delegation behavior.
That allows multiple differently composed agents to exist in one dsh process.
Current shipped preset IDs include:
standardcodeminimalcordis
The Web UI presents the last one as the Creator-oriented experience.
DeepSeek's four useful agent modes
The launch material describes four practical experiences.
Standard mode
This is the general coding agent. It exposes the normal toolbox for repository work: file access and search, shell, web capabilities when configured, planning, goals, skills, subagents and workflows.
Use it when you want a conventional full-featured coding agent.
Code mode
Code mode keeps Standard mode's capabilities but changes how tools are presented to the model.
Instead of requiring a model round trip for every small operation, the model can write a TypeScript program against a generated SDK and execute several operations through run_code.
Conceptually:
Normal tool calling
model -> tool A -> model -> tool B -> model -> tool C -> model
Code mode
model -> TypeScript program using A+B+C -> results -> model
The source configuration explicitly describes a sequence that could take five round trips being performed in one model-written program.
This can be valuable when a task needs predictable filtering, joining, aggregation or a set of dependent operations. It can reduce model round trips and prevent large intermediate results from repeatedly occupying model context.
It is not automatically superior. Programmatic orchestration is a poor fit when each observation should change the model's next decision, when actions require separate approvals, or when model-written orchestration code becomes harder to inspect than direct calls.
Minimal mode
Minimal mode is deliberately small: a fixed prompt plus persistent Bash and a str_replace_editor. It suppresses much of the richer runtime context and omits context compaction.
This makes it useful for benchmarking and harness research because fewer extra capabilities are influencing the model.
Minimal does not mean safest or best for daily use. It means less harness.
Creator mode / cordis preset
Creator mode is intended for people building the harness itself: inspecting the runtime, experimenting with Cordis plugins and creating custom agent presets.
It turns the harness into a development environment for new harness compositions.
The session log is a major design decision
One of the strongest parts of the architecture is the session log.
DeepSeek says the session log is the source of the context the model sees. Model history is derived from that log, and raw assistant chunks are retained for replay and UI fidelity.
The design rule is effectively:
If something becomes visible to the model, it must be reconstructable from the log.
That gives one durable event stream several jobs:
- session persistence;
- model-context reconstruction;
- transcripts;
- replay;
- resume;
- fork;
- telemetry; and
- debugging.
This matters when an agent fails after 30 tool calls. Without a coherent event history, “why did the model do that?” often degenerates into guesswork.
An event-sourced design does not guarantee reproducibility—external APIs and nondeterministic models can still change—but it gives the runtime a far better audit trail.
How one agent turn actually works
A useful way to understand any harness is to follow one turn.
At a simplified level, dsh:
- receives the user's task;
- builds the model-visible prompt and available tool schemas;
- sends a request through the selected LLM adapter;
- records the assistant output;
- intercepts requested tool calls;
- applies permission and sandbox rules;
- executes allowed tools;
- records tool outputs;
- decides whether another model step is required; and
- records the turn ending when the agent becomes quiescent.
A step is roughly one model request plus its tool work. A turn can contain multiple steps.
Plugins can attach around these lifecycle events, which is how behavior can be extended without replacing unrelated subsystems.
DeepSeek Harness is not tied to DeepSeek models
The name is easy to misread.
dsh ships with DeepSeek integration, but the provider layer is broader. The current model configuration guide documents:
- DeepSeek;
- Anthropic;
- OpenAI;
- AWS Bedrock;
- Google Vertex;
- Azure;
- Codex OAuth; and
- custom providers, including self-hosted or company OpenAI-compatible endpoints.
That means you can use DeepSeek Harness as the surrounding runtime while testing different model families.
Why this matters
Suppose you want to compare three models on the same repository task.
If each model is tested in its vendor's own agent product, you are comparing:
Model A + Harness A vs Model B + Harness B vs Model C + Harness C
With a sufficiently provider-neutral harness, you can hold more of the runtime constant:
Model A + dsh vs Model B + dsh vs Model C + dsh
That still is not a perfectly controlled experiment because models may expect different tool schemas, reasoning APIs or prompting styles, but it removes many confounding variables.
How to install DeepSeek Harness
Requirements
For the current JavaScript workspace, the repository declares Node:
^22.19.0 OR >=24.0.0
Because this is a developer preview, check the live README before automating installation around a pinned version.
Fastest route: Web UI
From a terminal in the project directory you want to work with:
npx @deepseek-ai/dsh web
The default Web UI is served at:
http://127.0.0.1:3080
You do not need to clone the repository just to evaluate the product.
Install and run from source
Use the source checkout when you want to inspect or modify plugins and presets:
git clone https://github.com/deepseek-ai/deepseek-harness.git
cd deepseek-harness
pnpm install
pnpm run build
pnpm dsh web
DeepSeek's root workspace currently uses pnpm.
First-run setup: a safe sequence
Do not begin by pointing a new autonomous agent at your home directory or a production repository.
Use this sequence instead.
1. Start inside a disposable repository
cd /path/to/disposable-project
npx @deepseek-ai/dsh web
The process uses its invoking directory as the default filesystem location, but a fresh Web UI still requires you to choose a workspace before composing a session.
2. Configure a model
Open:
Settings -> Models
For DeepSeek, add the API key.
For another supported catalog provider, choose Add provider and configure the appropriate credentials.
For a self-hosted model or company gateway, choose Add a custom provider and supply:
- provider ID;
- base URL;
- API protocol;
- credentials; and
- model information.
The provider ID should be chosen carefully because saved sessions and credential references use it.
3. Select the workspace
Choose only the repository you intend the agent to access.
4. Keep the safe permission preset
The current default permission preset is designed around:
workspace-write + ask
That means the runtime confines writes to the workspace and uses approval policy for actions that require escalation.
Do not switch to danger-full-access merely to eliminate prompts.
5. Begin with a read-only task
For example:
Summarize this repository. Identify the five files most important to the
authentication flow. Do not modify anything.
Inspect the trace and tool behavior before asking for mutations.
6. Run a contained implementation task
For example:
Find the root cause of the failing unit test, make the smallest correct fix,
run the focused tests, then run the relevant broader validation. Do not commit
or push anything.
This tests the loop, file editing, shell execution and verification without authorizing an external write.
Credentials: where keys are stored
DeepSeek documents model keys entered through the UI as write-only from the page. The literal secret is stored in:
$DSH_HOME/.credentials.yaml
Settings keep a credential reference, while the UI receives a redacted descriptor after saving.
That is useful secret separation, but it is not a complete security model.
A model does not need to “see the API key in Settings” to cause damage if a tool has enough filesystem, shell or network authority. Real security comes from the combination of:
- filesystem confinement;
- process sandboxing;
- network policy;
- narrow workspaces;
- approvals;
- credential isolation; and
- trustworthy plugins.
Privacy: local-first is not the same as fully offline
DeepSeek's Data Processing Statement describes Harness as local-first: session material, runtime state and configured credentials are primarily kept on the machine rather than making DeepSeek's cloud the mandatory storage layer. That is useful for self-hosting, but it does not mean every dsh session is private or offline.
Data can still leave the machine whenever the composed harness uses an external service, including:
- a cloud model provider;
- web search or retrieval;
- an MCP server;
- a third-party plugin; or
- a company gateway that logs requests.
So assess privacy at the whole harness graph, not only at the main LLM endpoint. A localhost model plus a cloud search plugin is not an offline agent. Session logs also deserve protection because they can contain prompts, source code, tool output, paths and operational context.
Sandboxing and permission presets
Agent security is one of the places where harness engineering directly changes what models can safely do.
DeepSeek Harness currently exposes two default high-level permission presets.
| Preset | Sandbox | Approval policy | Practical meaning |
|---|---|---|---|
workspace-write |
Workspace-write confinement | Ask | Normal default; work inside workspace, request approval for escalation |
danger-full-access |
No confinement | Never ask | Maximum autonomy and maximum local blast radius |
The sandbox layer itself recognizes:
read-only;workspace-write; anddanger-full-access.
The local sandbox has OS-specific implementations:
- Linux: bubblewrap / Landlock paths;
- macOS: Seatbelt;
- Windows: ACL restricted-token backend.
An important detail is that DeepSeek's sandbox mode governs filesystem effects. Do not interpret “workspace-write” as a universal network-isolation guarantee.
One-shot approvals
DeepSeek's approval service is designed around narrow decisions such as:
- allowed once;
- rejected;
- cancelled; or
- unavailable.
If no approval handler exists when approval is required, the design fails closed rather than silently proceeding.
That is preferable to an “approval requested, therefore assume yes” architecture.
Known developer-preview security limitations
DeepSeek's own repository is unusually candid about several boundaries that are still evolving. These are important if you are evaluating dsh as infrastructure rather than merely running a local demo.
web_fetch is intentionally disabled in shipped presets
The HTTP fetch provider currently implements URL validation, size/time limits and same-origin redirect restrictions, but its own README states that private-network / SSRF protection is deferred.
It does not yet block private, loopback, link-local or other non-public destinations after DNS resolution. DeepSeek therefore calls the provider an SSRF primitive when deployed where it can reach sensitive internal targets.
The shipped base composition responds correctly to that limitation: web_fetch is disabled and no HTTP fetch provider is mounted by default. Web search is a separate capability and remains available through the configured search provider.
Do not enable the raw fetch provider inside a network that can reach metadata endpoints, internal admin panels or other sensitive services unless you add an appropriate network boundary.
The Web UI is deliberately local
The Web server itself has no general TLS/authentication layer. The shipped CLI currently rejects:
dsh web --host 0.0.0.0
because all-interface exposure would turn a local agent capable of shell execution into a network-reachable remote-code-execution surface.
Keep the Web UI on loopback unless you have deliberately engineered a hardened deployment boundary. A reverse proxy is not automatically sufficient if the underlying API trust and authentication model has not been designed for that topology.
“Sandboxed” must be read precisely
The current local sandbox modes primarily describe filesystem effects. If your threat model includes data exfiltration or hostile code, also reason about:
- outbound network access;
- environment variables;
- inherited credentials;
- sockets and local services;
- container/VM boundaries; and
- what a plugin executes in the host process.
This is a general agent-security lesson: approval policy, filesystem confinement and network containment are different controls.
Security rules that matter in practice
Even a well-designed harness is not a safe default for every repository.
Use these rules:
- Treat repository content as potentially hostile. README files, issue text, build output and downloaded pages can contain prompt injection.
- Keep secrets outside the agent workspace.
- Use least privilege. Start with read-only or workspace-write.
- Do not use
danger-full-accessfor convenience. - Separate local authority from external authority. Editing a test file and pushing to a production branch are different risk classes.
- Review plugin provenance. “Everything is a plugin” creates flexibility but also creates a software supply-chain surface.
- Pin versions for serious evaluation. Preview behavior can change.
- Inspect the effective runtime. Use
--dump-configinstead of assuming a config file is the whole composition. - Keep production credentials out of test agents.
- Verify network controls separately. Filesystem confinement does not automatically mean egress confinement.
Anthropic's and Cursor's 2026 engineering reports on their own coding-agent sandboxes make the broader lesson clear: better containment can increase useful autonomy because the agent is allowed to work freely inside a hard boundary rather than interrupting the user for every ordinary command.
Run DeepSeek Harness without the Web UI
For automation, the headless profile is more interesting than the browser.
The current CLI supports a one-shot task such as:
dsh --profile headless "run the tests and explain the failures"
The headless runner creates a fresh persisted session, submits the task, waits for the agent to settle, flushes the session and prints the final assistant text.
The shipped headless composition opens no HTTP server or browser client.
That makes it suitable for:
- local scripts;
- repeatable evaluation;
- batch repository analysis;
- controlled automation jobs; and
- integrations where a browser UI would only add overhead.
Do not equate “headless” with “safe for unattended production.” The permission and execution environment still need to be engineered for the action being automated.
Python SDK: using dsh as infrastructure
DeepSeek also publishes a Python SDK route.
The current guide requires Python 3.10+ and shows:
python -m pip install deepseek-harness-sdk
The repository includes a runnable minimal JSON-RPC example that can be given:
- a workspace;
- a session root;
- a session ID; and
- a task.
Its session directory receives a JSONL log containing assembled model requests and tool calls.
This is the route to consider when you want the harness inside another application rather than as a standalone Web UI.
Real-world use cases
DeepSeek Harness is broader than “AI that writes code.” Its value is strongest when the environment itself needs to be configurable.
1. Repository engineering agent
Typical task:
Investigate why the build fails on Windows but passes on Linux. Reproduce it,
identify the smallest root cause, patch it, run the focused tests and report
remaining risk.
Harness features involved:
- repository search;
- file editor;
- shell;
- iterative loop;
- test feedback;
- session trace;
- sandbox;
- approval boundary.
2. Model bake-off under one runtime
Use Minimal or a controlled custom preset to test several models with the same:
- tool schema;
- prompt;
- shell;
- editor;
- timeout;
- workspace;
- stopping policy.
This is far more meaningful than comparing marketing benchmark numbers from unrelated products.
3. Internal enterprise agent
A company could build plugins for:
- internal source search;
- ticket systems;
- knowledge bases;
- deployment status;
- databases;
- proprietary scanners.
Then expose only the approved capabilities to a particular preset.
The challenge is governance: every new tool creates authority that must be scoped, logged and reviewed.
4. Security-review or incident-analysis assistant
A controlled agent can combine:
- repository inspection;
- logs;
- static-analysis output;
- web/research sources;
- shell tools;
- a dedicated read-only or isolated environment.
This can help with triage and evidence correlation. It should not be given production credentials merely because the task is defensive.
5. Headless maintenance automation
Examples:
- dependency inventory;
- changelog generation;
- stale-code analysis;
- documentation consistency checks;
- test-failure summaries.
A headless harness can perform these tasks on a schedule or from another orchestrator.
6. Harness research and plugin development
Creator mode is specifically valuable here.
A harness engineer can experiment with:
- alternative tool schemas;
- different loop logic;
- a remote sandbox;
- new context policies;
- custom subagent backends;
- alternative web search providers;
- verification plugins.
This is the area where dsh is most differentiated from polished closed-source coding products.
How harness design changes the same model
There are several direct mechanisms.
1. Tool descriptions shape decisions
A model can only choose tools intelligently if names, arguments, constraints and return formats are clear.
A poorly documented shell tool can cause:
- repeated failed commands;
- unnecessary privilege requests;
- huge outputs;
- destructive behavior.
Cursor publicly described improving its shell tool feedback so the model understood when a sandbox caused a failure and when escalation was appropriate. That is a pure harness change, yet it improved agent recovery.
2. Context selection controls attention
Dumping an entire repository into the prompt is not “more context engineering.” It can bury the useful evidence.
A harness decides:
- which files to retrieve;
- whether to summarize old tool output;
- what session state to preserve;
- when to compact;
- which instructions must remain stable.
Good context management increases the probability that the next model call sees the right evidence, not merely more tokens.
3. Verification creates actionable feedback
A model that writes a patch and immediately stops is much weaker operationally than the same model in a loop that runs:
test -> inspect failure -> revise -> test again
The test does not make the model smarter. It gives the model ground truth about its last action.
4. Stopping policy changes reliability and cost
Stop too early and the task remains unverified.
Stop too late and the agent burns tokens, repeats searches or damages a working solution.
A harness therefore needs a notion of:
- completion;
- pending tool work;
- failure;
- maximum turns;
- cancellation; and
- quiescence.
5. Sandboxing changes autonomy
Without containment, every shell command may deserve a human approval.
With a strong boundary, the harness can allow more autonomous exploration inside a disposable environment.
This can improve both safety and throughput.
6. Tool presentation changes token economics
DeepSeek's Code mode is a concrete example.
If the model can run several deterministic operations inside one generated program, it may avoid multiple model round trips and avoid serializing every intermediate result back into the context.
That can be a major efficiency gain for the right workflow.
7. Session design changes long-horizon work
Long-running agents fail when they forget earlier decisions, lose the state of a partial task or repeatedly rediscover the same facts.
A durable session log plus explicit resume/fork behavior gives the runtime a better foundation for long tasks.
DeepSeek Harness vs the main competitors
No single ranking is correct because these systems optimize for different goals.
| Product | Source model | Model flexibility | Runtime extensibility | Isolation / permissions | Best fit |
|---|---|---|---|---|---|
DeepSeek Harness (dsh) |
Open source, MIT | DeepSeek plus Anthropic/OpenAI/catalog/custom compatible providers | Very high; model, tools, loop, sessions, sandbox, UI and more are plugin-composed | Built-in sandbox modes + approvals; OS-specific local backends | Harness developers, controlled experiments, custom agent platforms |
| Claude Code | Closed product; Anthropic has open-sourced some supporting components such as sandbox runtime | Claude models through Anthropic-supported deployment routes | Strong user extensibility through MCP, skills, hooks, subagents and SDK, but core product is not a general replace-everything plugin kernel | Mature permissions, OS sandboxing, network/filesystem controls and auto mode | Polished Claude-first coding workflows |
| OpenAI Codex CLI | Open source, Apache-2.0 | OpenAI/Codex-first | Skills, MCP, app-server/SDK surfaces and configuration; less centered on replacing every internal subsystem | Built-in sandbox and approval model | Users who want an open local Codex client tightly integrated with OpenAI's coding stack |
| Gemini CLI | Open source, Apache-2.0 | Gemini-first | MCP, extensions, custom commands, hooks, subagents and skills | Confirmation policy plus optional Docker/Podman/Seatbelt and tool sandboxing | Google/Gemini users wanting an extensible terminal agent |
| OpenCode | Open source, MIT | Strong multi-provider orientation | Custom tools, MCP, agents, skills, plugins and client/server architecture | Fine-grained allow/ask/deny; shell still has host-user authority unless separately isolated | Provider-agnostic terminal coding with a mature open ecosystem |
| Cline | Open source, Apache-2.0 | Broad multi-provider + local model options | IDE/CLI/SDK, MCP, hooks/plugins, browser and tool approvals | Human-in-the-loop approvals; execution authority depends on configured environment | IDE-centric open harness and teams embedding a mature coding-agent runtime |
| LangChain Deep Agents / dcode | Open source, MIT | Model-agnostic | Harness profiles, middleware, tools, memory, skills, subagents and deployment surfaces | Execution backend is configurable; production isolation depends on chosen sandbox/deployment | Teams wanting model-specific harness tuning and a programmable open agent platform |
| OpenHands SDK / apps | Open source, MIT | Multi-provider | Composable SDK, tools, workspaces, server and applications | Local or remote/containerized workspaces with sandbox-oriented deployment options | Building and deploying software agents as an application platform |
| Cursor Agent | Closed-source product | Multi-model product with model-specific tuning | Rules, skills, plugins/MCP and editor integrations; core harness remains vendor-managed | Cross-platform local agent sandbox plus approval/run modes | Best-in-class integrated IDE workflow rather than harness research |
Where DeepSeek Harness is strongest
dsh has the clearest advantage when the question is:
“Can I replace or recombine the runtime itself?”
Its architecture is explicitly built around that goal.
Model adapter? Plugin.
Agent loop? Plugin.
Sandbox backend? Plugin.
Tool presentation? Plugin.
Session behavior? Plugin.
UI? Plugin.
That is more structurally open than products whose core loop is vendor-controlled but whose edges can be extended.
Where Claude Code is stronger
Claude Code is a much more mature finished coding product.
Anthropic has spent substantial engineering effort on:
- permissions;
- sandboxing;
- prompt-injection defenses;
- subagents;
- hooks;
- SDK use;
- long-running workflows; and
- high-polish developer UX.
If your goal is “use Claude to ship code today,” Claude Code is usually the more direct choice.
If your goal is “replace the loop, model provider or execution substrate and study the resulting agent,” DeepSeek Harness is conceptually more open.
Where Codex CLI is stronger
Codex CLI is open source and optimized as the local interface to OpenAI's coding-agent stack. It provides a direct path from ChatGPT/Codex access to repository work and has a strong sandbox/approval model.
Its center of gravity is using Codex effectively, not turning every internal runtime component into a generic plugin seam.
Where Gemini CLI is stronger
Gemini CLI has a mature extension system around:
- MCP;
- commands;
- hooks;
- skills;
- subagents; and
- sandbox options.
Its model path is naturally Gemini-centric, while dsh is explicitly experimenting with a more provider-neutral model layer.
Where OpenCode is stronger
OpenCode is a strong choice when the primary requirement is open-source, multi-provider coding with a polished terminal workflow.
Its current permission system can apply fine-grained allow, ask and deny policies by action and resource.
A crucial security distinction: its shell runs with the host user's filesystem, process and network authority. Permission rules are useful control logic, but they should not be confused with a hard OS/container isolation boundary. For high-risk autonomy, put it inside a separate VM/container or equivalent confinement.
Where Cline is stronger
Cline has a longer history as an end-user coding agent and a strong IDE-first workflow. Its current open-source stack supports many cloud providers plus OpenAI-compatible and local-model routes, and it combines editor diffs, terminal execution, browser work, MCP and explicit approvals. Its SDK also makes the same general coding-agent capabilities available to developers building their own integrations.
Choose Cline when you want an established open IDE agent with a programmable SDK. Choose dsh when the main goal is to replace deeper runtime subsystems such as the loop, session machinery or execution providers through one composition model.
Where LangChain Deep Agents / dcode is stronger
Deep Agents is one of the closest open architectural comparisons because it is explicitly model-agnostic and treats harness behavior as something developers should tune. Its harness profiles can vary system prompts, tool descriptions, middleware and subagent behavior per provider or model, and dcode packages that stack as an open terminal coding agent.
It currently has a stronger story around data-driven model-specific harness tuning: LangChain publishes controlled results showing that the same model can improve substantially when prompts, tools and middleware are changed. dsh is more radical about making nearly every runtime capability a Cordis plugin, while Deep Agents is more immediately connected to the broader LangChain/LangSmith application and deployment ecosystem.
Where OpenHands is stronger
OpenHands has evolved into a broader software-agent platform with a composable SDK, workspace abstractions, remote execution and applications.
It is arguably the closest direct architectural competitor to dsh.
Both emphasize modularity and event/state discipline. OpenHands has more history as an application and deployment ecosystem; dsh pushes especially hard on runtime recomposition through Cordis and plugin seams.
Harnesses vs LangGraph, AutoGen and CrewAI
These are worth comparing, but they answer a different first question.
LangGraph / AutoGen / CrewAI style systems:
“How should I define and orchestrate my own agent application or multi-agent workflow?”
DeepSeek Harness / Claude Code / Codex / OpenCode style systems:
“How should a model actually operate inside a real working environment right now?”
There is overlap. You can build a coding agent with a framework, and you can embed or automate a harness. But choosing between them usually depends on whether you need a library for authoring application logic or a ready execution runtime with tools, sessions and safety boundaries.
DeepSeek Harness strengths
1. Architectural openness
The loop and infrastructure are not treated as sacred internals.
2. Provider flexibility
You can test models beyond DeepSeek and configure custom compatible endpoints.
3. Event-sourced session design
Model-visible state, replay and persistence have a coherent source of truth.
4. Multiple runtime shapes
Web UI, headless execution and Python SDK cover interactive, automated and embedded use cases.
5. Purpose-built modes
Standard, Code, Minimal and Creator expose genuinely different harness strategies rather than cosmetic personas.
6. Security is a subsystem, not an afterthought
Sandbox policy, approvals and permission presets have explicit architecture.
7. MIT license
The core project is permissively licensed for experimentation and integration.
Weaknesses and risks
1. It is a developer preview
This is the largest issue. DeepSeek explicitly warns about compatibility-breaking changes.
Do not build an expensive internal platform assuming current package names and configuration are stable APIs.
2. Plugin flexibility becomes operational complexity
An “everything is replaceable” architecture is powerful for harness engineers and potentially overwhelming for normal users.
More composition layers mean more possible misconfiguration.
3. Plugin supply-chain risk
A plugin is executable code with whatever authority the composed runtime grants it.
A third-party plugin can become a more serious security boundary than the model itself.
4. Cross-provider support does not mean equal model quality
Different models are trained and tuned around different tool schemas, reasoning APIs and interaction patterns.
A provider-neutral harness can normalize the environment, but the best prompt and tool presentation for one model may not be best for another.
5. Trace logs can contain sensitive material
Auditability is valuable, but durable session logs may contain:
- source code;
- prompts;
- tool output;
- file contents;
- command output; and
- operational context.
Treat session storage as sensitive data.
6. Local sandboxing is not the same as total containment
The current sandbox documentation is explicit that sandbox modes govern filesystem effects. A high-assurance environment may also need:
- network egress restrictions;
- disposable containers or VMs;
- secret brokers;
- scoped credentials;
- CPU/memory limits; and
- separate external-action proxies.
7. Fast-moving documentation
The repository is moving quickly enough that current master, npm release tags and documentation may not always describe exactly the same packaged version.
Pin what you evaluate.
Should you switch from another coding agent?
Use this decision guide.
Choose DeepSeek Harness if
- you are building agent infrastructure, not only consuming an agent;
- you want to replace the model provider, tools, loop, sandbox or session behavior;
- you need a controlled model-harness benchmark environment;
- you are researching Code mode or programmatic tool orchestration;
- you want a permissively licensed base for custom agent products; or
- you accept preview churn.
Choose Claude Code if
- you primarily want a polished Claude-native coding agent;
- mature permissions and sandbox UX matter more than replacing core internals;
- you value Anthropic's model-specific harness tuning.
Choose Codex CLI if
- you want a strong open-source local coding client for OpenAI/Codex;
- your workflow already centers on ChatGPT/Codex;
- you prefer the vendor's model-harness co-design over broad provider neutrality.
Choose Gemini CLI if
- you are invested in Gemini and Google tooling;
- Google Search grounding, extensions and MCP are central;
- you want a mature open-source CLI rather than a harness laboratory.
Choose OpenCode if
- you want a polished open-source, multi-provider terminal coding agent;
- broad provider choice matters;
- you are willing to add stronger external isolation when needed.
Choose Cline if
- you want an open-source agent deeply integrated into the IDE;
- broad provider and local-model support matter;
- you want an SDK around an already mature coding-agent workflow.
Choose Deep Agents / dcode if
- you want an open, model-agnostic harness with explicit per-model tuning profiles;
- LangGraph/LangSmith observability and deployment fit your stack;
- middleware, memory, skills and subagent composition matter more than replacing every runtime service.
Choose OpenHands if
- you want an open software-agent SDK/application platform;
- remote sandboxed execution and deployment are important;
- you are building services around agents rather than only a terminal workflow.
Choose Cursor if
- IDE-native experience and multi-model convenience matter most;
- you want the vendor to tune the harness for each model;
- you do not need the core runtime to be open or replaceable.
A practical benchmark before choosing any harness
Do not choose from one demo.
Create 10–30 tasks drawn from your actual work, for example:
- diagnose a failing test without edits;
- fix a small bug;
- add a feature touching three modules;
- perform a security review;
- update docs after an API change;
- refactor with behavior-preserving tests;
- analyze a large log;
- recover from a deliberately broken tool call.
For each model-harness pair, record:
| Metric | Why it matters |
|---|---|
| Success / failure | Basic capability |
| Tests actually run | Whether completion is evidence-backed |
| Tokens per success | Cost efficiency |
| Wall time | User throughput |
| Tool calls | Orchestration efficiency |
| Repeated calls | Loop quality |
| Human approvals | Supervision burden |
| Out-of-scope actions | Alignment with intent |
| Recovery from denial/failure | Harness feedback quality |
| Quality of final summary | Usability and audit value |
Only then decide whether “better model” or “better harness” is responsible for the result you care about.
Frequently asked questions
Is DeepSeek Harness a model?
No. It is the runtime around models. A model supplies reasoning and generation; dsh supplies the environment, tools, loop, sessions, permissions and other agent infrastructure.
Do I need a DeepSeek model?
No. DeepSeek is the first-party path, but current documentation includes Anthropic, OpenAI and custom provider routes among others.
Can I use a local model?
Potentially, yes. The model server must expose an API route compatible with a configured provider/adapter. A custom provider can point to a self-hosted server.
Local inference does not automatically make the entire harness offline. Audit web tools, telemetry, plugins and any other external providers separately.
Is DeepSeek Harness free?
The source is MIT-licensed. That does not make model inference, hosted search providers or other third-party services free.
Is it production ready?
DeepSeek says no such thing; it calls the project a developer preview and warns of breaking changes. Production use requires your own stability, security, upgrade and operational testing.
Is a harness the same as an agent?
Not quite. The model plus the harness produces the operational agent. In everyday speech, people often call the whole product “the agent.”
Is MCP a competitor to dsh?
No. MCP is an interoperability protocol that a harness can use to reach external tools and data. dsh is the larger runtime.
Can a better harness make a weak model as good as a frontier model?
Not reliably. A harness can eliminate avoidable execution failures and use feedback more efficiently, but it cannot consistently overcome missing underlying reasoning or knowledge.
Why would Code mode be faster or cheaper?
When several tool operations can be expressed as one bounded program, the model may avoid multiple round trips and large repeated intermediate outputs. That can reduce latency and token use. It needs to be benchmarked on the actual workload.
Why keep Minimal mode if Standard has more features?
Because “more harness” can confound model evaluation. Minimal mode deliberately removes capabilities so researchers can observe a smaller model-plus-tools system.
What is the biggest security mistake?
Granting broad host and credential access because the model appeared reliable in a few successful sessions. Agent failures are not only hallucinated text; with tools they become real actions.
The deeper lesson: models are becoming only half the product
The old mental model was:
Pick the smartest model.
The more accurate agent-era question is:
Pick the model, harness and authority boundary that work together for this task.
Model training determines what reasoning is possible.
Harness engineering determines what the model can observe, what it can act on, how it gets feedback, what it remembers, how it recovers, when it stops and how much damage it can cause.
DeepSeek Harness is important because it exposes those choices instead of hiding them behind one fixed agent product.
Its “everything is a plugin” architecture is not automatically better for every user. For most developers who simply want a coding assistant, a mature closed or opinionated agent may be easier. For researchers and engineers who want to build the runtime itself, however, dsh is one of the most explicit attempts yet to make the entire agent stack swappable and inspectable.
That is the real reason to pay attention to it.
Sources and further reading
This guide prioritizes first-party documentation and primary research. Product behavior is time-sensitive; facts below were re-checked on August 21, 2026.
DeepSeek Harness
- DeepSeek Harness developer preview
- DeepSeek Harness Data Processing Statement
- DeepSeek Harness repository and README
- Architecture: Cordis, profiles, session log and capability seams
- Web UI guide
- Model provider configuration
- Python SDK guide
- Permission presets
- Process sandbox
- HTTP fetch provider and its SSRF limitation
- Shipped base composition:
web_fetchdisabled - Web server security and loopback posture
- CLI/headless reference
- Code preset composition
- Minimal preset composition
Harness engineering and research
- OpenAI: Harness engineering — leveraging Codex in an agent-first world
- Anthropic: Effective harnesses for long-running agents
- LangChain: Improving Deep Agents with harness engineering
- LangChain: Tuning Deep Agents to Work Well with Different Models
- The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- What makes a harness a harness