Connector benchmark · Round two
810 runs27 tasks3 independent judgesAugust 2026
One connector finished ~2.5× more of what was asked.
We ran the same battery of real reporting and diagnostic questions against Skai's Reporting MCP and against a publisher's own MCP, ten times each, and had three independent judges grade every answer. Skai cleared the completeness bar about two and a half times as often, on 5.2× fewer tokens. Where a publisher's MCP does answer, it is usually accurate and clearly written — the gap is in follow-through, not correctness.
Figure 1
What one connector reaches that five cannot
A cross-publisher answer prints like a colour plate: one channel per plate, all of them registered against the same keyline, and the sheet is only whole once every plate lands. 13 of the 25 questions in this battery need more than one plate. This is the only finding in the report that doesn't narrow as publisher MCPs improve.
A single publisher's own MCP
Owns one plate. However fast or cheap it becomes, the other three were never its data.
The sheet never finishes printing.
Skai's Reporting MCP
Every plate registers against the same keyline, so the blended answer can be printed at all.
One question, one answer, every channel in it.
Schematic, not measured — the figure shows which data exists on each side, and carries no invented values. The counts behind it are real: of the 25 questions in the battery, 13 span more than one publisher and were reworded into a same-question, single-publisher version before the single-publisher side was scored on them.
Figure 2
What each question cost, both sides
One shared scale, one row per question. The vertical rule is the floor: the publisher MCP never answered anything for less than 143k tokens, because a 100+-tool schema goes into every call whether the question needs it or not. Skai's cheapest question cost 6k.
Performance
What changed in my account in the last few days?
Tool calls
1 vs. 1
Seconds to answer
6s vs. 14s
Tokens per question
5,620 vs. 143,352
Runs behind each median
10 per side
Waste
Data quality
Triage
Year over year
Each bar is the median of 10 runs of that question, on both sides, in one round of testing. The vertical rule is the cheapest question the publisher MCP answered all round — nothing it did cost less than that, because a 100+-tool schema is carried into every call whether the question needs it or not. Skai's side has no such floor: its cheapest question in the battery cost 6k. The 4 multi-turn flows in the same battery are held out of this chart — they chain several questions into one session, so they aren't comparable to a single ask.
Figure 3
The same cost, projected past what we measured
Skai's per-query cost barely moves as accounts are added. Publisher MCPs carry their full schema on every call, so the gap widens with scale rather than narrowing. Measured head-to-head cost per query runs 1.5× against small read-only publisher schemas and up to 8× against a large one, corroborated in this round at 4.1×–5× the tokens per query.
Projected monthly cost against accounts in scope, for Skai's Reporting MCP and for calling publisher MCPs directly. At 90 accounts the projection marks $39.62 a month for Skai's Reporting MCP against $394.12 for calling publisher MCPs directly. Both curves are modeled beyond a measured 2-account baseline.
- Skai's scaling exponent — measured
- The same cross-account question asked against 1 account and against 10 under one login: 58,607 vs. 74,602 tokens per query. A 27% rise for ten times the accounts, because a cross-account question costs about one extra call however many accounts are in scope.
- ~0.10
- The publisher exponent — modeled
- An engineering assumption, not a regression. We ran the identical 1-versus-10 test on that side and could not use the result: a prompt-wording anomaly on the single-account run, plus up to 48% day-to-day variance on repeated identical queries. It stays as the prior modeled figure rather than being replaced by an unreliable measurement.
- 0.6
- Where the curve is anchored
- Both sides start from directly observed per-query averages at 2 accounts. Everything to the right of that is the model, not a measurement.
- 2 accounts
Plotted at a reference footprint of 3 publishers × 30 accounts and 300 reporting questions a month — 90 accounts in scope, marking $39.62 against $394.12 a month. Run it against your own footprint on the other tab.
How it was run
Two rounds. Round one: five reporting workflows, run twice each — once through Skai's Reporting MCP, once calling a publisher's MCP directly — against publishers whose own MCPs ship small, read-only tool schemas. Round two: 27 real reporting and diagnostic tasks, 10 runs each, in two account-scoping configurations of a publisher MCP shipping 100+ tools across reporting, account management, and campaign management, graded independently by three judges — one automated, two manual — over roughly 2,300 ballots.
“Real, data-backed answer” means judged completeness above the rubric floor: not a bare refusal, not an unresolved clarifying question standing in for the whole answer, not an empty or errored response. Fig. 2's bars are the median of the 10 runs of each question on each side, from the publisher MCP's stronger account-scoping mode rather than its weaker one. Pricing throughout is Claude Sonnet 5 at $2.00/MTok in and $10.00/MTok out, cache tokens excluded.
Back to the interactive view to run the cost model against your own footprint.
Read the method, then run it yourself
Every number above came out of one connector answering real questions about real accounts. Setup is an OAuth flow and a URL.