For anyone tracking athletic performance with a smartwatch or chest strap, the assumption that all metrics are equally reliable deserves serious scrutiny. As wearable adoption accelerates across amateur and elite team sports alike, a clearer picture of where these devices succeed — and where they quietly mislead — has real implications for training decisions, injury prevention, and recovery management.

This systematic review screened PubMed and Scopus for studies published between 2015 and 2025, ultimately analyzing eleven studies that compared wearable-device outputs against gold-standard measurement methods in team-sport athletes. Heart rate emerged as the most reliably captured metric, maintaining high validity across device types and particularly under controlled laboratory conditions. Respiratory frequency also showed strong agreement with criterion measures when purpose-built respiratory monitoring devices were used. However, energy expenditure estimation was a consistent weak point: wearables demonstrated substantial variability and a pattern of systematic underestimation during high-intensity and intermittent efforts — precisely the movement signatures that define team sports like soccer, basketball, and rugby. VO2max estimation showed mixed results, with accuracy varying considerably by device model and testing protocol.

These findings align with a growing body of validation literature suggesting that optical and accelerometer-based sensors perform well for simple, rhythmic physiological signals but struggle with the metabolic complexity of stop-start, multi-directional sport. The eleven-study sample is a meaningful limitation — it restricts the statistical power to detect device- or sport-specific patterns — and laboratory validity does not always translate to field conditions. Practically, this means caloric output data from consumer wearables should be treated as rough estimates rather than precise training inputs, while heart rate data remains the most actionable and trustworthy metric. The review is incremental rather than paradigm-shifting, but it consolidates a fragmented evidence base into a useful decision framework for coaches and performance staff weighing which wearable outputs merit operational trust.