benchmark106 tasks$17.51 total spend

Four budget models. One harness.
The whole benchmark.

We ran OpenDesktop through all 106 tasks of Harness-Bench with four of the paper's exact model backends. One beats the best published harness average. The whole campaign billed $17.51.

The result

The Harness-Bench paper publishes full-suite completion scores for six agent harnesses. Here is that table with OpenDesktop added:

completion · published harnesses vs this run · full suite
harnesscompletiontokens/task
OpenDesktop × GLM 5.1 82.6% 108k
NanoBot 81.6% 5.0k
Hermes 80.4% 139.7k
OpenDesktop (4-backend avg) 80.1% 113k
Moltis 78.4% 134.9k
NullClaw 75.9% 175.1k
ZeroClaw 69.9% 133.2k
OpenClaw 60.0% 74.0k

With GLM 5.1, OpenDesktop completes 82.6% of Harness-Bench. That is above the best average any published harness posted. Averaged across all four backends we ran, we land at 80.1%, right behind the second best, and there is no frontier model anywhere in our mix.

What we ran

Harness-Bench (Qihoo 360, 2026) measures agent harnesses, not models. 106 knowledge-worker tasks: email triage, office documents, incident analysis, data cleaning, injection defense. A programmatic oracle grades the files the agent leaves on disk.

We drove the same OpenDesktop binary users download, headlessly, through its local HTTP API. No special benchmark build, no per-task tuning. The benchmark's own proxy measured every token.

Four of the paper's eight model backends. All 106 tasks each. One run per cell, oracle scoring only. Everything billed $17.51, and every dollar figure on this page comes from that invoice.

The four models

Same suite, same harness, four backends from the paper's own lineup. What each one scored, consumed, and cost:

full suite · 4 backends × 106 tasks
#modeloutcometokens/taskcost/taskwhole run
1 GLM 5.1 82.6% 108k $0.067 $7.13
2 Kimi k2.5 79.5% 92k $0.034 $3.58
3 Qwen 3.6-plus 79.2% 113k $0.047 $4.94
4 DeepSeek v4-flash 79.0% 138k $0.018 $1.86

costs are each model's actual OpenRouter invoice. full per-task data (CSV)

The most expensive model tops the board at seven cents per completed task. The other three finish within half a point of each other, at two to five cents.

score vs billed cost per task (log scale)
78% 80% 82% 84% $0.01 $0.02 $0.05 $0.10 billed cost per task (log scale) GLM 5.1: 82.6% at $0.067/task GLM 5.1 Kimi k2.5: 79.5% at $0.034/task Kimi k2.5 Qwen 3.6-plus: 79.2% at $0.047/task Qwen 3.6-plus DeepSeek v4-flash: 79.0% at $0.018/task DeepSeek v4-flash

That flatness is the finding. The harness does enough of the work (workspace anchoring, sandboxed execution, document skills, recovery from malformed tool calls) that swapping models moves the score by single points. It is also the bet OpenDesktop makes: the default model is a cheap one, because on real desk work a pricier model mostly buys a bigger bill.

show all 106 per-task scores
per-task oracle scores · 106 × 4
taskflashkimiglmqwen
001-file 1.001.001.001.00
002-exec 1.001.001.001.00
003-browser 1.001.001.001.00
004-meeting-summary 0.890.891.000.89
005-email-triage 1.001.001.001.00
006-access-bilibili 1.001.001.001.00
007-session-memory 1.001.001.000.75
008-image-recognize 1.001.001.001.00
009-git-pr-merge 0.250.250.250.25
010-office-docs 1.001.001.001.00
011-code-debug 0.860.860.860.86
012-doc-synthesis 0.750.750.750.75
013-image-edit 1.001.001.001.00
014-task-decomposition 0.940.860.690.69
015-security-injection-defense 0.701.000.700.70
016-code-repair-pytest 0.901.001.001.00
017-db-doc-consistency 0.970.900.970.97
018-provider-failover-audit 0.750.860.920.93
019-incident-runbook-synthesis 0.820.810.820.82
020-archive-checksum 0.711.001.000.71
021-batch-rename-transform 0.820.820.910.71
022-local-rest-api-summary 0.650.650.650.61
023-web-form-extraction 0.840.740.840.84
024-calendar-scheduling-conflict 0.890.890.780.89
025-meeting-action-tracker 0.940.890.940.50
026-ppt-brief-generation 0.900.800.900.80
027-contract-summary-risk 0.710.790.930.71
028-email-thread-merge 0.640.640.730.82
029-expense-packet-review 0.770.001.001.00
030-word-revision-plan 0.880.880.881.00
031-cross-doc-citation-check 0.440.780.670.44
032-customer-followup-draft 0.860.711.001.00
033-offline-knowledge-qa 0.770.770.771.00
034-evidence-matrix-claims 0.880.920.920.92
035-conflicting-source-resolution 0.840.840.840.84
036-citation-consistency-audit 0.780.780.780.78
037-policy-clause-retrieval 0.740.740.740.74
038-research-brief-synthesis 0.970.920.820.90
039-repo-architecture-map 0.840.950.870.90
040-test-coverage-fill 1.001.000.430.43
041-frontend-state-bug 0.600.600.600.59
042-api-schema-migration 0.960.740.740.74
043-db-migration-safety 0.920.870.990.98
044-ci-config-repair 1.001.001.001.00
045-dependency-upgrade-compat 0.680.750.740.60
046-performance-regression 1.001.001.000.70
047-code-review-risk-report 0.640.740.620.74
048-release-note-changelog 0.590.740.990.96
049-excel-like-cleaning 1.000.921.001.00
050-multitable-join-analysis 1.001.001.001.00
051-sql-query-report 0.691.001.001.00
052-metric-definition-audit 0.690.691.001.00
053-anomalous-transaction-detect 0.931.000.980.98
054-budget-variance-analysis 1.000.691.001.00
055-funnel-dropoff-analysis 0.930.930.920.93
056-inventory-forecast 0.690.690.690.69
057-interruption-resume 0.810.811.000.81
058-multiday-project-state 0.920.920.920.92
059-event-update-replan 1.000.361.001.00
060-task-cancellation-cleanup 0.821.000.491.00
061-periodic-status-rollup 0.731.001.001.00
062-k8s-config-audit 0.860.890.910.86
063-alert-dedup-noise 0.970.910.971.00
064-service-dependency-triage 0.840.840.840.82
065-capacity-planning 0.670.830.790.81
066-rollback-readiness 0.820.870.920.89
067-canary-release-check 0.900.931.000.96
068-product-launch-ops 0.850.920.850.92
069-legal-compliance-review 0.951.001.001.00
070-hr-resume-screening 0.700.880.700.88
071-ecommerce-support-routing 0.830.830.830.83
072-logistics-delay-response 0.880.680.880.88
073-research-repro-package 0.500.600.401.00
074-education-grading-feedback 0.860.660.860.66
075-platform-appeal-review 0.590.860.860.76
076-medical-admin-claim-check 0.830.780.890.78
077-archive-manifest-defense 0.480.450.830.28
078-local-api-cursor-retry-ledger 0.000.560.590.59
079-smallfile-batch-reject-ledger 0.490.610.530.02
080-schema-roundtrip-conversion 0.740.910.910.91
081-local-html-dom-form-extract 0.850.700.800.90
082-compose-config-repair 1.001.000.950.95
083-monorepo-interface-repair 0.951.000.950.23
084-js-state-type-bug 1.000.600.920.57
085-flaky-test-root-cause 0.990.180.180.18
086-sql-migration-preflight-rollback 0.960.670.950.63
087-cli-parser-bug-tests 0.620.820.290.52
088-api-contract-mock-client-compat 0.400.500.400.50
089-ab-test-caveat-analysis 0.951.000.950.95
090-timeseries-anomaly-attribution 0.540.630.560.53
091-financial-close-reconciliation 0.480.410.840.38
092-schema-drift-audit 0.340.500.500.38
093-jsonl-sessionization-analysis 0.400.410.590.59
094-metric-definition-migration-diff 0.780.970.780.83
095-policy-version-conflict-resolution 0.740.690.690.69
096-offline-knowledge-qa-insufficient-evidence 0.650.650.650.65
097-research-claims-batch-evidence-audit 0.720.720.720.92
098-three-source-decision-record-synthesis 0.530.450.690.41
099-privacy-dsar-intake-review 0.910.601.000.91
100-financial-kyc-admin-check 0.700.700.600.70
101-marketing-sensitive-commitment-review 0.720.900.900.90
102-internal-doc-retrieval-injection-defense 0.600.600.600.60
103-policy-update-replan-diff 0.840.800.931.00
104-async-ops-window-rollup 0.861.000.870.87
105-partial-batch-resume-ledger 0.670.640.700.60
106-release-approval-gate-plan 0.510.950.910.93

some rows are identical across all four models. Those are rubric and environment effects, not model skill; we show them rather than pruning them.

What this does and does not show

  • All 106 tasks ran. One (a git workflow) is structurally impossible for our sandbox image today and scored near-floor for every backend; it is included in every mean, not excluded.
  • Outcome scoring only. The paper also grades process and security with an LLM rubric; we used the programmatic oracles only. Full grading would lower our score, as it did every published harness's.
  • One run per cell. Single-task scores carry noise; the raw data marks where all four models score identically.
  • We ran it ourselves. Which is why the per-task scores, token counts, and adapter are published, and the benchmark is third-party and open.

Reproduce it

The benchmark is public: Qihoo360/harness-bench (we ran commit 1025086a). Our entire integration is one 280-line adapter file plus a config entry per model. The per-task scores are raw oracle output, unedited. If you run it and get different numbers, we want to hear about it.

The harness is the product.

OpenDesktop is the agent on your desktop: your real files, a local sandbox, a visible per-request cost, and a default model this benchmark says you should not feel bad about.

Download OpenDesktop How it works