Every score on this site traces back to the same testing framework. Here's how it works.

The testing period

Machines live in our test kitchen for a minimum of three weeks of daily use โ€” because week-one impressions lie. Heat-up claims get verified on day one; workflow annoyances, cleaning burdens and consistency drift reveal themselves by week three.

What we measure

Espresso machines: brew temperature stability (thermocouple and Scace-style measurement across 10 consecutive shots), shot timing variance, steam power (time to texture 150ml to 140ยฐF), heat-up time, and noise (decibel meter at 1m). Grinders: particle distribution (sieve testing), retention (weighed in vs out), dosing consistency across 10 doses, and noise. Brewers: brew-water temperature at the shower head through a full cycle, brew time, and carafe heat retention. Frothers and accessories: task-specific protocols described within each review.

What we taste

Instruments can't taste balance. Each machine brews the same reference coffees (a medium washed blend, a light single origin, a dark roast) and our panel evaluates blind against reference machines wherever possible. Milk texture is assessed by pour behavior and latte art performance.

Scoring

Four sub-scores (0โ€“10): Performance, Ease of Use, Value, and Cleaning & Maintenance. The overall star rating (out of 5) weighs these against the product's price class โ€” a $100 product scoring 9 means best-in-class at $100, not better than a $2,000 machine.

Long-term follow-up

We keep reference units in rotation and update reviews when durability problems emerge, firmware changes behavior, or better rivals reset the category. A review's "last updated" reflects genuine re-evaluation, not a date-stamp refresh.

What we don't do

No pay-for-review, no score previews for manufacturers, no untested "roundups" assembled from spec sheets. If we haven't used it, we don't score it. Our funding model is described in the affiliate disclosure.