Abstract
Over the past decades, advances in computer vision and deep learning have made it possible to capture human movement in three dimensions using only video. This progress has opened the door to scalable, accessible analyses of human motion outside of the lab; however, the practical application of these methods to biomechanics remains limited by challenges in accuracy, generalization, and biomechanical interpretability. This dissertation aims to bridge that gap by validating, scaling, and translating 3D pose estimation techniques into meaningful biomechanical contexts.
The first study investigates the accuracy of modern pose estimation systems for ergonomic posture assessment through a controlled experiment focused on comparing monocular, stereo, and depth-assisted camera setups against optical motion capture. Across 46 participants and 3,832 joint-angle observations, deep learning-based monocular models achieved the lowest error (geometric mean RMSE near 13°), outperforming both stereo and depth-assisted configurations and validating their potential for large-scale workplace analysis without specialized hardware.
The second study addresses biomechanical interpretability by mapping 3D pose data to standardized joint angles following International Society of Biomechanics conventions, using a lightweight learned model trained on a large-scale athletic motion capture dataset. By linking estimated poses to interpretable joint kinematics such as flexion/extension and abduction/adduction, this work moves beyond spatial accuracy toward true biomechanical meaning. Benchmarking three state-of-the-art lifting models through this mapping reveals a paradox: the model with the lowest positional error (MPJPE) produces the highest joint-angle error, showing that MPJPE by itself is insufficient for evaluating biomechanical fidelity. These results motivate the use of joint kinematics as constraints in the training of 2D-to-3D lifting models.
The final study develops an automated framework to extract and analyze professional tennis serve biomechanics from broadcast footage, combining detection, pose estimation, 3D lifting, and temporal segmentation into a single pipeline. Applied to 5,966 serves from 109 professional players, the resulting dataset reproduces known kinetic-chain biomechanics and reveals individualized “kinematic fingerprints,” with a classifier identifying players from joint-angle trajectories alone at 99.2% accuracy. The resulting dataset captures authentic, high-speed athletic motion and reproduces laboratory-tested results, demonstrating that markerless systems can yield laboratory-quality insights at scale.
Together, these studies validate and extend 3D pose estimation for real-world biomechanics, advancing the study of human movement across occupational health and sports performance and laying the groundwork for future use of markerless motion capture.