Chimera 2.0: the most capable marketing model, benchmarked against frontier models and marketing tools.

Announcements · Model Card

Introducing Chimera 2.0

The most capable marketing model in the world, and the first built to run campaigns end to end, not just write about them.

Published
June 29, 2026
Category
Model launch
Read
6 min
Powers
Ares Operator

Today we are releasing Chimera 2.0, the model behind Ares, our autonomous marketing operator. Where general-purpose frontier models are trained to be good at everything, Chimera 2.0 is trained to be exceptional at one thing: turning a marketing objective into booked revenue, with no human stitching the steps together.

It sets a new state of the art on every marketing benchmark we tested, including MARK-bench, our end-to-end evaluation of an agent's ability to research a market, build the audience, generate the creative, launch the campaign, and optimize toward conversions, autonomously.

MARK-bench: Agentic marketing (end-to-end)
% of campaigns reaching the conversion target autonomously · higher is better
Chimera
Chimera 2.071.2%
Frontier models
Opus 4.853.1%
GPT-5.249.5%
Gemini 3 Pro46.8%
Llama 431.4%
Marketing tools
HubSpot Breeze27.3%
Jasper22.6%
Copy.ai18.9%

The gap is widest exactly where it matters. Frontier models write excellent ad copy in isolation, but degrade sharply once a task requires holding a budget, a compliance boundary, and a multi-step funnel in mind at once. Point-solution marketing tools never attempt the agentic loop at all: they generate an asset and stop.

Built to operate, not just write

Chimera 2.0 was post-trained on real campaign trajectories: spend decisions, audience iterations, compliance checks, and the messy reality of leads who do not reply on the first touch. The result is a model that is competitive with the best frontier systems on raw copy, and decisively ahead on everything that happens after the copy is written.

FunnelBench

Funnel architecture quality (0–100)
Chimera 2.078.5
Opus 4.858.9
GPT-5.255.0
Gemini 3 Pro52.3
HubSpot Breeze38.7

Outreach-bench

Cold SMS / email reply rate (%)
Chimera 2.041.7%
Opus 4.830.5%
GPT-5.228.3%
Gemini 3 Pro26.9%
HubSpot Breeze23.4%

BrandVoice

Voice consistency across 50 outputs (%)
Chimera 2.092.3%
Opus 4.885.2%
GPT-5.281.5%
Gemini 3 Pro79.0%
Jasper74.6%

Compliance

TCPA / A2P / DND correctness (%)
Chimera 2.099.1%
Opus 4.891.2%
GPT-5.288.4%
Gemini 3 Pro85.7%
HubSpot Breeze81.2%
Full results

Across all eight benchmarks, Chimera 2.0 ranks first, leading the strongest frontier model (Claude Opus 4.8) by an average of 14.6 points and the best dedicated marketing tool by 31.2 points.

MARK-bench

% of campaigns reaching the conversion target autonomously · higher is better
Chimera 2.071.2%
Opus 4.853.1%
GPT-5.249.5%
Gemini 3 Pro46.8%
Llama 431.4%
HubSpot Breeze27.3%
Jasper22.6%
Copy.ai18.9%

AdCopy-Eval

conversion-weighted copy quality
Chimera 2.064.8
Opus 4.862.7
GPT-5.261.2
Gemini 3 Pro58.4
Jasper51.3
Copy.ai49.7
HubSpot Breeze46.2
Llama 444.1

FunnelBench

Funnel architecture quality (0–100)
Chimera 2.078.5
Opus 4.858.9
GPT-5.255.0
Gemini 3 Pro52.3
HubSpot Breeze38.7
Llama 433.0
Jasper24.1
Copy.ai20.4

BrandVoice

Voice consistency across 50 outputs (%)
Chimera 2.092.3%
Opus 4.885.2%
GPT-5.281.5%
Gemini 3 Pro79.0%
Jasper74.6%
Copy.ai71.2%
HubSpot Breeze70.1%
Llama 468.4%

Outreach-bench

Cold SMS / email reply rate (%)
Chimera 2.041.7%
Opus 4.830.5%
GPT-5.228.3%
Gemini 3 Pro26.9%
HubSpot Breeze23.4%
Jasper21.0%
Copy.ai19.8%
Llama 419.2%

Compliance

TCPA / A2P / DND correctness (%)
Chimera 2.099.1%
Opus 4.891.2%
GPT-5.288.4%
Gemini 3 Pro85.7%
HubSpot Breeze81.2%
Llama 476.3%
Jasper64.0%
Copy.ai60.5%

Local-SEO

GBP + citations
Chimera 2.088.0
Gemini 3 Pro73.1
Opus 4.871.9
GPT-5.270.4
Llama 458.0
HubSpot Breeze55.4
Jasper41.2
Copy.ai38.9

ROAS-Sim

return index, normalized
Chimera 2.03.9×
Opus 4.82.6×
GPT-5.22.4×
Gemini 3 Pro2.3×
HubSpot Breeze1.9×
Llama 41.6×
Jasper1.4×
Copy.ai1.3×

Available today

Chimera 2.0 powers every Ares operator account at no additional cost. API access is rolling out to design partners this quarter.

Start with Ares →
Modelchimera-2.0
Context1M tokens
ChannelsSMS · Email · Ads · GBP
Operating modeAutonomous + approval-gated
PricingIncluded in Ares

Methodology. All benchmarks were run on a held-out suite of 1,200 real-world marketing tasks across home-services, e-commerce, and local-business verticals. Agentic benchmarks (MARK-bench, FunnelBench, Outreach-bench) measure autonomous completion of a full task trajectory; static benchmarks are scored by a panel of human marketing reviewers blind to model identity.

Frontier models were evaluated through their public APIs with an identical marketing tool harness. Marketing tools were evaluated using their native agent or generation features. ROAS-Sim figures are simulated against historical spend data and are directional, not guaranteed outcomes.

Comparison figures are illustrative and intended for demonstration. Model names are trademarks of their respective owners.

Put an operator on the payroll

Ares starts at $300 a month and works every hour of it.