← All case studies
NewApplied AIAgentic Systems0 to 1

Carbii — An Independent Agent for the Dealers the Big Tools Don’t Reach

Independent Builder · AI Product Management Capstone, Product Faculty (8-week course)

01
Carbii sidebar showing an AI-drafted reply for a marketplace lead

The challenge

Independent dealers get leads from every marketplace, dumped unstructured into one Gmail inbox, with no tooling at all

What I did

Over an 8-week AI PM capstone at Product Faculty, built a 4-agent system, solo, that drafts and screens replies inside Gmail

Result

A 594-run, cross-provider model comparison cut pipeline latency 40%

"A 594-run, cross-provider model evaluation cut pipeline latency 40% — rigor that paid for itself in the metric that actually matters to a dealer: speed."

Problem

I'd already seen this problem up close — I shipped the AI-response feature described in Case Study 2 for mobile.de's largest, most professionalised dealers. But that tool only exists for dealers already inside a Lead Management System. 71% of independent dealers have no LMS at all — just Gmail, and marketplace leads arriving with no structure, no tracking, and no help. The average response time is 9.2 hours; 82% of buyers expect a reply within 10 minutes. The dealers who need this most are the ones no enterprise tool is built for.

Approach

Rather than another dashboard dealers would have to adopt, I built Carbii to meet them exactly where they already are — a Chrome extension that activates inside Gmail itself. Underneath, it's a 4-agent pipeline, not a single model call: a Guard agent screens every incoming message for manipulation before anything else touches it and fails closed on any doubt; a Conversation Agent drafts the reply and classifies intent signals; an independent Integrity Critic re-checks every draft against guardrails before a dealer ever sees it; a Bulk Reply agent batches genuinely similar leads.

I treated model selection as an empirical question, not a default: a 594-run comparison (9 models × 66 eval cases, cross-provider LLM judging so no provider grades its own family) drove the final production model choices — and directly cut the synchronous reply pipeline's latency by 40%. Nothing sends without the dealer's explicit approval.

Outcomes

3
Marketplaces covered — AutoScout24, mobile.de, Kleinanzeigen
66
Eval cases spanning all 4 agents, cross-provider judged
594
Model comparison runs behind the production model choices
-40%
Cut in synchronous reply pipeline latency

Key insight

Shipping the enterprise version taught me the pattern works. Building the independent version taught me the harder lesson — for the dealers with nothing, the bar isn't a better dashboard, it's zero behaviour change. Evaluation rigor wasn't academic overhead here — it's what actually found the 40% speed-up.

← Back to portfolioNext →
AI-Generated Lead Responses
Get in touchLinkedIn