
Handling manual assembly tasks is a crucial part of everyday life. Especially with surgical instruments, proficient handling is critical for flawless workflows, efficiency, and overall patient safety. Markerless object pose estimation and augmented reality (AR) visualization can guide users through the assembly process. There is a lack of understanding when combining state-of-the-art object tracking with different types of AR assembly visualizations. Our work introduces a comparison between tabletop in-situ, FoV-fixed, and world-fixed visual guidance techniques. We use a benchmark task with step-by-step assembly guidance via gaze control and reproducible 3D-printed assemblies of varying complexity and shape (proxies for surgical instruments). Our within-subject user study with 32 participants (16 scrub nurses and 16 laypersons) revealed that AR assembly guidance can match and partially outperform paper-based instructions regarding usability, user experience, and workload. Objective metrics (times and errors) were largely comparable to the paper baseline. We identified that world-fixed and FoV-fixed animations tend to overall outperform in-situ animations, which by design require sequential visual processing and whose augmentations’ precision is most affected by pose estimation inaccuracies. We also contrast characteristics of expert and layperson feedback. All participants commended the gaze-driven testbed interface for step-by-step AR instructions, but also revealed needs for further optimization and individualization. Our insights on visualization placement characteristics using pose estimation and involving a spectrum of users and assemblies open a new perspective in the field and aim to foster future work towards pose estimation-driven AR assembly guidance