In 2025, METR paid experienced open-source developers $150 an hour to do real work, randomly assigning each task to AI-allowed or AI-disallowed. Before starting, the developers predicted AI would cut their completion time by 24%. Afterward, they estimated it had cut it by 20%. The stopwatch said they were 19% slower.
Sit with that. These are technical people who measure things for a living, and they were wrong about their own productivity by roughly 40 percentage points — in the flattering direction.
Now ask yourself how confident you actually are that the AI powered executive assistant you pay for every month is producing anything at all. Not how it feels. What it produced.
Try the Easy Button™ Sales Intelligence System Free for 30 Days
No risk. No long-term contract. We’ll prove ROI before you spend a dollar — or you walk away. Also includes:
- →Free Market Snapshot — See how many ideal-fit companies exist in your target market
- →1,000 Free Verified Prospect Contacts — Real decision-makers matching your Ideal Customer Profile. Yours to keep.
- →30-Day Proof of ROI — Walk away if it doesn’t work
Time saved is not a business outcome
Salesforce research from 2023 found sales reps spend under 30% of their week actually selling. Every automation vendor on earth quotes that number at you. Almost none of them ask the second question, which is the only one that matters:
When you free up an hour, where does the hour go?
In most small and mid-sized companies, it goes right back into the pile. More email. More Slack. More checking on something. Time you reclaim without redirecting is time you donate back to entropy. You feel less busy and your pipeline looks identical.
That is the structural flaw in how the AI executive assistant category was designed. It was built around friction reduction — fewer clicks, fewer calendar conflicts, shorter email threads. Friction reduction is real, and it is almost completely invisible on a P&L. Which is why, when the CFO asks what the tool did, nobody has an answer that survives follow-up questions.
What if the assistant had a quota?
Our thesis is that the job description is wrong, not the technology.
An AI powered executive assistant should not be evaluated on how much time it hands back. It should be evaluated on the quality of the decisions it hands you. Same tooling, different scoreboard.
Instead of manage my calendar, the job becomes: every morning, name the five people most likely to buy from me this week, and show your work.
That is a testable job. You can grade it. Run the five names it gives you against five names pulled at random from the same database, call both sets, and compare booked meetings after four weeks. If the assistant loses, fire it. You cannot do that with time saved, because nobody has ever audited a time-saved claim in the history of software.
The architecture we would build
This is a proposed workflow, not a case study. Here is the shape of it:
SIGNALS IN NORMALIZE SCORE COMPRESS ACT ----------- --------- ----- -------- --- Email opens/clicks -> Site page views -> One contact -> Weighted -> Top 5 by -> Human CRM stage changes -> record, one recency + delta this call, Ad engagement -> identity intent + week, with today Form + chat events -> graph fit the reason
Four things make this work, and none of them are the model:
1. Identity resolution. If Sarah opening an email and Sarah reading your pricing page are two unrelated rows, you have data, not signal. This is unglamorous CRM automation and data hygiene work, and it is where most builds die.
2. Weighted scoring, not point-counting. An email open is not a pricing-page visit. A pricing-page visit at 11pm from a director-level title at a company inside your ICP is not the same as a competitor snooping. Our Easy Button lead scoring system exists to make that difference explicit rather than intuitive.
3. Delta, not absolute score. The highest-scoring contact in your database is often someone who has been warm for two years and will never buy. What you want is who moved this week. Movement is the buying signal. Altitude is not.
4. Compression to a human-sized list. Watchtower exists for exactly this: thousands of behavioral events collapsed into a short list a human can act on before lunch. Give a salesperson 200 scored leads and you have given them a spreadsheet. Give them five names and a sentence each and you have given them a morning.
The economics, honestly
Illustrative example — run your own numbers. Say a rep makes 25 outbound calls a day and books 1 meeting per 25 dials from a cold, unsorted list. That is 1 meeting a day, 20 a month.
Now assume prioritization only doubles the connect-to-meeting rate — a modest assumption, not a promised result. Same 25 dials, 2 meetings a day, 40 a month. At a $12,000 average contract value and a 20% close rate, that difference is roughly $48,000 a month in new bookings from the same headcount and the same phone.
The point is not the number. The point is the leverage location. You are not buying more activity. You are buying better ordering of activity you were already paying for. A human executive assistant who saves you six hours a week costs $4,000 a month and cannot see your website traffic. Software that reorders your call list can, and it costs a fraction of that.
What the evidence actually says
The honest read on AI assistants is more mixed than the vendor decks suggest, and you should know it before you spend.
Gartner predicted in June 2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, polling over 3,400 organizations. The cited causes were escalating costs, unclear business value, and inadequate risk controls — not model failure. Translation: the projects die because nobody defined the job.
The strongest positive evidence has a skill gradient. Brynjolfsson, Li and Raymond studied 5,179 customer support agents and found a 14% average productivity lift, rising to 34% for novices and close to zero for the most experienced workers. AI assistance compressed the gap between the median performer and the best one. It did not make the best one better.
And self-reported gains are not trustworthy. METR surveyed 349 technical workers in 2026 and found a median self-reported 1.4x to 2x change in the value of their work — while explicitly listing reasons to be skeptical of that magnitude. When the researchers doing the survey warn you about their own headline number, believe them.
Where this gets uncomfortable
Three admissions, because you will hit all three.
If you are the experienced operator, this helps your team more than it helps you. That NBER skill gradient cuts directly against the pitch. If you have twenty years of pattern recognition and you already know which six accounts matter, a scoring engine is going to tell you things you knew. Buy it for your junior AE and your newest hire. That is where the 34% lives.
If your database has no engagement, there is nothing to route. This is the failure we find most often. A company has 6,000 contacts and 140 of them have ever opened anything. Scoring that produces confident nonsense — an ordered list of noise. If that is you, you do not have a prioritization problem, you have a top-of-funnel problem, and you should fix it with TAM mining and outbound plus a consistent content and distribution engine before you buy a single scoring rule. We would rather find that out in week two than sell you a system you cannot feed.
And the science is moving. METR themselves withdrew confidence in their own slowdown finding in February 2026, reporting that selection effects — developers refusing to work without AI — made the newer data unreliable, and that speedups now look likely. That is what intellectual honesty looks like, and it means our thesis deserves the same treatment. Measure your own system. Do not take ours, or theirs, on faith.
The playbook — build a working version in a week
You do not need a platform purchase to test this. Steal the sequence:
Day 1 — Pick three signals, not thirty. Email click, pricing or high-intent page view, and CRM stage change. Three is enough to prove or kill the thesis. Thirty is a six-month project you will abandon.
Day 2 — Fix identity. Make sure email address is the join key across your email platform, your site analytics, and your CRM. If it is not, stop here and fix it. Everything downstream is worthless otherwise.
Day 3 — Write the scoring rules on paper first. Literally paper. Pricing page = 15. Email click = 5. Second visit in 7 days = 20. Title match to ICP = 10. Argue about the weights out loud with whoever sells. You will learn more in that argument than in any tool evaluation.
Day 4 — Build the delivery, not the dashboard. A daily 7am email with five names, five reasons, and five phone numbers. No dashboard. Nobody logs into dashboards. The output of an AI powered executive assistant should arrive where the human already is.
Day 5 — Instrument the feedback loop. Track which of the five got called and which converted. Feed outcomes back into the weights monthly. Without this step you have built a random number generator with good branding.
Run it for four weeks against a random-list control. If the scored list does not beat random, your weights are wrong or your signal is thin — and either answer is worth more than another month of guessing. If you want someone to own that build and the argument that comes with it, that is what a fractional CMO engagement is for.
The takeaway
Your prospects are already talking. Build something that listens. Connect the systems. Track the behavior. Score meaningful engagement. Compress thousands of activities into a manageable action list. Then put your salespeople where they belong: talking to the humans most likely to care.
An AI powered executive assistant that manages your calendar makes your week smoother. One that routes your decisions makes your quarter. Pick the second job description.
Want to build it yourself? Steal the framework above — it is the whole thing. Want us to bolt it onto the stack you are already paying for? See what this costs or work with Jeremy directly for 90 days.
