ARR Doesn't Mean What It Used To
VCs are no longer trusting ARR at face value - here's exactly how founders are gaming metrics, what investors now watch for instead, and why the traditional SaaS benchmarking playbook broke in the AI era.
Read Original Summary used for search
TLDR
• ARR has fractured into multiple definitions: some founders annualize daily revenue (×365), others use quarterly×4, and many conflate GMV with actual recurring revenue
• CARR (contracted ARR) worries investors more than inflated ARR - large drop-offs between signed contracts and realized revenue are common, especially in AI training data businesses
• Gross margin manipulation is rampant: AI companies push FDE/integration costs below the line into OpEx to show "70%+ GM" when those costs actually scale linearly with revenue
• The old playbook (benchmark growth/NRR/burn against comps) doesn't work anymore - AI's extreme growth rates force qualitative judgment over quantitative heuristics
• Physical AI has the same problem wearing different clothes: autonomy rates measured in controlled pilots and non-binding LOIs presented as pipeline
In Detail
The piece compiles direct quotes from VCs and partners at major funds revealing how founders are gaming metrics and what red flags investors now watch for. ARR has quietly stopped meaning one thing - investors report seeing it calculated as quarterly revenue ×4, daily revenue ×365, or conflated with GMV. The consensus: founders should define their ARR calculation upfront rather than let investors guess.
CARR (contracted/committed ARR) worries investors even more because conversion rates are unpredictable. A crossover fund partner notes the gap between contracted and realized revenue can be "quite large" - customers change their minds, don't onboard, or face unpredictable implementation periods. For AI training data companies specifically, eight-figure SOWs signal demand but recognition depends on milestone delivery and the vendor's ability to scale increasingly complex work as model capabilities evolve. Gross margin is equally suspect: AI-native companies with FDE motions push integration costs into OpEx to show clean margins, but those costs scale linearly with revenue growth. NRR gets inflated by including POC-to-contract conversions that aren't comparable to traditional expansion metrics.
The deeper shift is that traditional SaaS benchmarking broke. Pre-AI, investors could compare growth/NRR/GRR/burn across companies and use those as investment heuristics. AI's extreme growth rates and hybrid pricing models (recurring + consumption) make quantitative benchmarking unreliable - it's now about qualitative judgment on durability. Physical AI shows the same pattern: autonomy rates measured in controlled pilots with engineers nearby, and non-binding LOIs presented as revenue pipeline. The through-line: numbers generated under ideal conditions get quoted as if those conditions are permanent. Investors want to know what happens after 90 days in production, not week one of the pilot.