Open source · 36 tasks · 4 agents

A benchmark for AI productivity agents

36 tasks. 12 personas. 5 scoring dimensions. Every result is a real deliverable in a real file format.

Lemon Supercomputer

Lemon Supercomputer is an AI productivity assistant for Mac. It runs a team of autonomous agents that handle real work: drafting documents, building spreadsheets, researching competitors, generating designs, pulling emails and calendar events. It connects to Gmail, Slack, Google Calendar, and multiple other integrations.

The benchmark

There are plenty of evaluations for the models themselves, and those are useful for measuring raw reasoning and knowledge. But the model is only one piece. What matters to the end user is the harness and architecture around it: how the agent plans tasks, calls tools, handles integrations, manages files, and delivers a finished result. That is what this benchmark tests.

36 tasks, each something a real professional would ask an AI assistant to do on a workday. Can the agent draft your sales strategy, build your dashboard, pull your calendar, and deliver a polished brand kit?

We created 36 tasks across 12 professional personas: roles like Personal Assistant, Sales Manager, Data Wizard, Designer, and Research Assistant. Each persona has three tasks at increasing difficulty: simple, moderate, and complex. Every task produces a real deliverable in a real file format. We tested four agents on all 36 and scored each output across five weighted dimensions.

Results

Lemon Supercomputer
8.5~3.1 min
Perplexity
7.6~8-10 min
Claude Co-work
7.4~47 sec
Manus
6.5~9.3 min

By complexity

Each persona has one task at each complexity level, giving 12 tasks per tier.

Simple tasks are single-step requests: summarize my calendar, generate a logo, find 10 companies matching a criteria. They test whether the agent can produce a correct output in the right format.

Moderate tasks require pulling from multiple sources or producing more structured output: a morning briefing combining calendar, email, Slack, and weather, or a competitive analysis with cited sources across three products. They test coordination and depth.

Complex tasks chain together research, synthesis, and polished deliverables: build an interactive HTML dashboard from raw CSV data, or pull email and Slack context about a specific investor and turn it into a tailored pitch deck. They test end-to-end workflow execution.

Scores tend to converge on simple tasks and diverge on complex ones.

TierLemon SCPerplexityCo-workManus
Simple87.176.5
Moderate8.47.77.36.7
Complex987.86.4

By persona

A persona is a professional role the agent is asked to play. Personal Assistant pulls your calendar and emails. Sales Manager preps discovery calls and builds pipeline reports. Designer generates logos and brand kits. There are 12 in total, each with three tasks at increasing difficulty.

PersonaLemon SCPerplexityCo-workManus
Personal Assistant8.77.76.35
Writer8.57.77.77
Comms Manager8.4776.3
Data Wizard8.4887
Marketing Expert87.37.36.7
Sales Manager8.5887.3
Lead Generation8.7886
Product Manager8.57.77.77.3
Travel Agent8.47.37.37
Designer8.67.36.77
Presentation Builder8.47.37.35.3
Research Assistant8.77.77.36.3

Example tasks

Every prompt is copied verbatim into the agent with no modifications. Here are six from across the difficulty range.

TASK 1Personal Assistant / Simple

What's on my calendar today? Pull my real calendar events and summarize them.

TASK 7Lead Generation / Simple

Find me 10 B2B SaaS companies in Austin, Texas with 50-200 employees that have raised Series A or B funding. Use real company data.

TASK 18Sales Manager / Moderate

I have a discovery call with the VP of Engineering at Rippling tomorrow. Prep me with company research, their recent product launches, likely pain points, and 10 discovery questions using the SPIN framework.

TASK 22Designer / Moderate

Create 3 Instagram ad creative images for an AI meeting notes app called 'MeetingMind'. Each should have a headline, subtext, and visual design. Generate actual image files, not text descriptions.

TASK 28Data Wizard / Complex

I have 6 months of sales data in the attached CSV. Build me an interactive dashboard I can open in a browser. Save as an HTML file.

TASK 35Presentation Builder / Complex

Research heylemon.ai and create a 15-slide investor pitch deck. Include market size with real data, competitive landscape, product differentiation, business model, and financial projections. Generate actual slides, not a text outline.

All 36 prompts are in the repository.

Scoring

Each task is scored across five dimensions on a 1 to 10 scale, then combined into a weighted composite.

30%
Output Quality
Accuracy, depth, usefulness
25%
Task Compliance
Right format, scope, deliverables
15%
Speed
Time to completion, scaled by tier (e.g. under 2 min for simple, under 10 min for complex)
15%
Integration & Tooling
Real APIs vs simulated data
15%
Polish & Presentation
Professional readiness

Full rubric with per-task breakdowns is in the repository.

Run it yourself

The benchmark is open source. Three paths depending on your agent type: a single mega-prompt for autonomous agents, a three-session split for chat-based agents, or a manual runbook for maximum control. The repo includes an auto-scorer that sends each artifact to Claude with the rubric and returns per-dimension scores.

View on GitHub

Try it yourself

See what your Mac
can really do.

Available on Mac. Free to download.

Download for Mac — freeFree download · macOS only · Set up in minutes