The result
The Harness-Bench paper publishes full-suite completion scores for six agent harnesses. Here is that table with OpenDesktop added:
| harness | completion | tokens/task |
|---|---|---|
| OpenDesktop × GLM 5.1 | 82.6% | 108k |
| NanoBot | 81.6% | 5.0k |
| Hermes | 80.4% | 139.7k |
| OpenDesktop (4-backend avg) | 80.1% | 113k |
| Moltis | 78.4% | 134.9k |
| NullClaw | 75.9% | 175.1k |
| ZeroClaw | 69.9% | 133.2k |
| OpenClaw | 60.0% | 74.0k |
With GLM 5.1, OpenDesktop completes 82.6% of Harness-Bench. That is above the best average any published harness posted. Averaged across all four backends we ran, we land at 80.1%, right behind the second best, and there is no frontier model anywhere in our mix.
What we ran
Harness-Bench (Qihoo 360, 2026) measures agent harnesses, not models. 106 knowledge-worker tasks: email triage, office documents, incident analysis, data cleaning, injection defense. A programmatic oracle grades the files the agent leaves on disk.
We drove the same OpenDesktop binary users download, headlessly, through its local HTTP API. No special benchmark build, no per-task tuning. The benchmark's own proxy measured every token.
Four of the paper's eight model backends. All 106 tasks each. One run per cell, oracle scoring only. Everything billed $17.51, and every dollar figure on this page comes from that invoice.
The four models
Same suite, same harness, four backends from the paper's own lineup. What each one scored, consumed, and cost:
| # | model | outcome | tokens/task | cost/task | whole run |
|---|---|---|---|---|---|
| 1 | GLM 5.1 | 82.6% | 108k | $0.067 | $7.13 |
| 2 | Kimi k2.5 | 79.5% | 92k | $0.034 | $3.58 |
| 3 | Qwen 3.6-plus | 79.2% | 113k | $0.047 | $4.94 |
| 4 | DeepSeek v4-flash | 79.0% | 138k | $0.018 | $1.86 |
costs are each model's actual OpenRouter invoice. full per-task data (CSV)
The most expensive model tops the board at seven cents per completed task. The other three finish within half a point of each other, at two to five cents.
That flatness is the finding. The harness does enough of the work (workspace anchoring, sandboxed execution, document skills, recovery from malformed tool calls) that swapping models moves the score by single points. It is also the bet OpenDesktop makes: the default model is a cheap one, because on real desk work a pricier model mostly buys a bigger bill.
show all 106 per-task scores
| task | flash | kimi | glm | qwen |
|---|---|---|---|---|
| 001-file | 1.00 | 1.00 | 1.00 | 1.00 |
| 002-exec | 1.00 | 1.00 | 1.00 | 1.00 |
| 003-browser | 1.00 | 1.00 | 1.00 | 1.00 |
| 004-meeting-summary | 0.89 | 0.89 | 1.00 | 0.89 |
| 005-email-triage | 1.00 | 1.00 | 1.00 | 1.00 |
| 006-access-bilibili | 1.00 | 1.00 | 1.00 | 1.00 |
| 007-session-memory | 1.00 | 1.00 | 1.00 | 0.75 |
| 008-image-recognize | 1.00 | 1.00 | 1.00 | 1.00 |
| 009-git-pr-merge | 0.25 | 0.25 | 0.25 | 0.25 |
| 010-office-docs | 1.00 | 1.00 | 1.00 | 1.00 |
| 011-code-debug | 0.86 | 0.86 | 0.86 | 0.86 |
| 012-doc-synthesis | 0.75 | 0.75 | 0.75 | 0.75 |
| 013-image-edit | 1.00 | 1.00 | 1.00 | 1.00 |
| 014-task-decomposition | 0.94 | 0.86 | 0.69 | 0.69 |
| 015-security-injection-defense | 0.70 | 1.00 | 0.70 | 0.70 |
| 016-code-repair-pytest | 0.90 | 1.00 | 1.00 | 1.00 |
| 017-db-doc-consistency | 0.97 | 0.90 | 0.97 | 0.97 |
| 018-provider-failover-audit | 0.75 | 0.86 | 0.92 | 0.93 |
| 019-incident-runbook-synthesis | 0.82 | 0.81 | 0.82 | 0.82 |
| 020-archive-checksum | 0.71 | 1.00 | 1.00 | 0.71 |
| 021-batch-rename-transform | 0.82 | 0.82 | 0.91 | 0.71 |
| 022-local-rest-api-summary | 0.65 | 0.65 | 0.65 | 0.61 |
| 023-web-form-extraction | 0.84 | 0.74 | 0.84 | 0.84 |
| 024-calendar-scheduling-conflict | 0.89 | 0.89 | 0.78 | 0.89 |
| 025-meeting-action-tracker | 0.94 | 0.89 | 0.94 | 0.50 |
| 026-ppt-brief-generation | 0.90 | 0.80 | 0.90 | 0.80 |
| 027-contract-summary-risk | 0.71 | 0.79 | 0.93 | 0.71 |
| 028-email-thread-merge | 0.64 | 0.64 | 0.73 | 0.82 |
| 029-expense-packet-review | 0.77 | 0.00 | 1.00 | 1.00 |
| 030-word-revision-plan | 0.88 | 0.88 | 0.88 | 1.00 |
| 031-cross-doc-citation-check | 0.44 | 0.78 | 0.67 | 0.44 |
| 032-customer-followup-draft | 0.86 | 0.71 | 1.00 | 1.00 |
| 033-offline-knowledge-qa | 0.77 | 0.77 | 0.77 | 1.00 |
| 034-evidence-matrix-claims | 0.88 | 0.92 | 0.92 | 0.92 |
| 035-conflicting-source-resolution | 0.84 | 0.84 | 0.84 | 0.84 |
| 036-citation-consistency-audit | 0.78 | 0.78 | 0.78 | 0.78 |
| 037-policy-clause-retrieval | 0.74 | 0.74 | 0.74 | 0.74 |
| 038-research-brief-synthesis | 0.97 | 0.92 | 0.82 | 0.90 |
| 039-repo-architecture-map | 0.84 | 0.95 | 0.87 | 0.90 |
| 040-test-coverage-fill | 1.00 | 1.00 | 0.43 | 0.43 |
| 041-frontend-state-bug | 0.60 | 0.60 | 0.60 | 0.59 |
| 042-api-schema-migration | 0.96 | 0.74 | 0.74 | 0.74 |
| 043-db-migration-safety | 0.92 | 0.87 | 0.99 | 0.98 |
| 044-ci-config-repair | 1.00 | 1.00 | 1.00 | 1.00 |
| 045-dependency-upgrade-compat | 0.68 | 0.75 | 0.74 | 0.60 |
| 046-performance-regression | 1.00 | 1.00 | 1.00 | 0.70 |
| 047-code-review-risk-report | 0.64 | 0.74 | 0.62 | 0.74 |
| 048-release-note-changelog | 0.59 | 0.74 | 0.99 | 0.96 |
| 049-excel-like-cleaning | 1.00 | 0.92 | 1.00 | 1.00 |
| 050-multitable-join-analysis | 1.00 | 1.00 | 1.00 | 1.00 |
| 051-sql-query-report | 0.69 | 1.00 | 1.00 | 1.00 |
| 052-metric-definition-audit | 0.69 | 0.69 | 1.00 | 1.00 |
| 053-anomalous-transaction-detect | 0.93 | 1.00 | 0.98 | 0.98 |
| 054-budget-variance-analysis | 1.00 | 0.69 | 1.00 | 1.00 |
| 055-funnel-dropoff-analysis | 0.93 | 0.93 | 0.92 | 0.93 |
| 056-inventory-forecast | 0.69 | 0.69 | 0.69 | 0.69 |
| 057-interruption-resume | 0.81 | 0.81 | 1.00 | 0.81 |
| 058-multiday-project-state | 0.92 | 0.92 | 0.92 | 0.92 |
| 059-event-update-replan | 1.00 | 0.36 | 1.00 | 1.00 |
| 060-task-cancellation-cleanup | 0.82 | 1.00 | 0.49 | 1.00 |
| 061-periodic-status-rollup | 0.73 | 1.00 | 1.00 | 1.00 |
| 062-k8s-config-audit | 0.86 | 0.89 | 0.91 | 0.86 |
| 063-alert-dedup-noise | 0.97 | 0.91 | 0.97 | 1.00 |
| 064-service-dependency-triage | 0.84 | 0.84 | 0.84 | 0.82 |
| 065-capacity-planning | 0.67 | 0.83 | 0.79 | 0.81 |
| 066-rollback-readiness | 0.82 | 0.87 | 0.92 | 0.89 |
| 067-canary-release-check | 0.90 | 0.93 | 1.00 | 0.96 |
| 068-product-launch-ops | 0.85 | 0.92 | 0.85 | 0.92 |
| 069-legal-compliance-review | 0.95 | 1.00 | 1.00 | 1.00 |
| 070-hr-resume-screening | 0.70 | 0.88 | 0.70 | 0.88 |
| 071-ecommerce-support-routing | 0.83 | 0.83 | 0.83 | 0.83 |
| 072-logistics-delay-response | 0.88 | 0.68 | 0.88 | 0.88 |
| 073-research-repro-package | 0.50 | 0.60 | 0.40 | 1.00 |
| 074-education-grading-feedback | 0.86 | 0.66 | 0.86 | 0.66 |
| 075-platform-appeal-review | 0.59 | 0.86 | 0.86 | 0.76 |
| 076-medical-admin-claim-check | 0.83 | 0.78 | 0.89 | 0.78 |
| 077-archive-manifest-defense | 0.48 | 0.45 | 0.83 | 0.28 |
| 078-local-api-cursor-retry-ledger | 0.00 | 0.56 | 0.59 | 0.59 |
| 079-smallfile-batch-reject-ledger | 0.49 | 0.61 | 0.53 | 0.02 |
| 080-schema-roundtrip-conversion | 0.74 | 0.91 | 0.91 | 0.91 |
| 081-local-html-dom-form-extract | 0.85 | 0.70 | 0.80 | 0.90 |
| 082-compose-config-repair | 1.00 | 1.00 | 0.95 | 0.95 |
| 083-monorepo-interface-repair | 0.95 | 1.00 | 0.95 | 0.23 |
| 084-js-state-type-bug | 1.00 | 0.60 | 0.92 | 0.57 |
| 085-flaky-test-root-cause | 0.99 | 0.18 | 0.18 | 0.18 |
| 086-sql-migration-preflight-rollback | 0.96 | 0.67 | 0.95 | 0.63 |
| 087-cli-parser-bug-tests | 0.62 | 0.82 | 0.29 | 0.52 |
| 088-api-contract-mock-client-compat | 0.40 | 0.50 | 0.40 | 0.50 |
| 089-ab-test-caveat-analysis | 0.95 | 1.00 | 0.95 | 0.95 |
| 090-timeseries-anomaly-attribution | 0.54 | 0.63 | 0.56 | 0.53 |
| 091-financial-close-reconciliation | 0.48 | 0.41 | 0.84 | 0.38 |
| 092-schema-drift-audit | 0.34 | 0.50 | 0.50 | 0.38 |
| 093-jsonl-sessionization-analysis | 0.40 | 0.41 | 0.59 | 0.59 |
| 094-metric-definition-migration-diff | 0.78 | 0.97 | 0.78 | 0.83 |
| 095-policy-version-conflict-resolution | 0.74 | 0.69 | 0.69 | 0.69 |
| 096-offline-knowledge-qa-insufficient-evidence | 0.65 | 0.65 | 0.65 | 0.65 |
| 097-research-claims-batch-evidence-audit | 0.72 | 0.72 | 0.72 | 0.92 |
| 098-three-source-decision-record-synthesis | 0.53 | 0.45 | 0.69 | 0.41 |
| 099-privacy-dsar-intake-review | 0.91 | 0.60 | 1.00 | 0.91 |
| 100-financial-kyc-admin-check | 0.70 | 0.70 | 0.60 | 0.70 |
| 101-marketing-sensitive-commitment-review | 0.72 | 0.90 | 0.90 | 0.90 |
| 102-internal-doc-retrieval-injection-defense | 0.60 | 0.60 | 0.60 | 0.60 |
| 103-policy-update-replan-diff | 0.84 | 0.80 | 0.93 | 1.00 |
| 104-async-ops-window-rollup | 0.86 | 1.00 | 0.87 | 0.87 |
| 105-partial-batch-resume-ledger | 0.67 | 0.64 | 0.70 | 0.60 |
| 106-release-approval-gate-plan | 0.51 | 0.95 | 0.91 | 0.93 |
some rows are identical across all four models. Those are rubric and environment effects, not model skill; we show them rather than pruning them.
What this does and does not show
- All 106 tasks ran. One (a git workflow) is structurally impossible for our sandbox image today and scored near-floor for every backend; it is included in every mean, not excluded.
- Outcome scoring only. The paper also grades process and security with an LLM rubric; we used the programmatic oracles only. Full grading would lower our score, as it did every published harness's.
- One run per cell. Single-task scores carry noise; the raw data marks where all four models score identically.
- We ran it ourselves. Which is why the per-task scores, token counts, and adapter are published, and the benchmark is third-party and open.
Reproduce it
The benchmark is public: Qihoo360/harness-bench (we ran commit 1025086a). Our entire integration is one 280-line adapter file plus a config entry per model. The per-task scores are raw oracle output, unedited. If you run it and get different numbers, we want to hear about it.
The harness is the product.
OpenDesktop is the agent on your desktop: your real files, a local sandbox, a visible per-request cost, and a default model this benchmark says you should not feel bad about.