Abstract
Modern vision-language models achieve strong predictive performance, yet their reasoning processes remain difficult to inspect, ground, and verify. This dissertation studies this problem through the lens of a structured intermediate: an intermediate representation that is explicit, grounded in visual evidence, and independently verifiable. We investigate how such structure can be enforced at four complementary levels: the learned representation, the inference procedure, the model computation, and the training objective. At the representation level, we show that concept bottleneck models can expose human-readable concepts without ensuring that those concepts faithfully track the input: visually imperceptible perturbations can substantially alter reported concepts while preserving the final prediction. We then develop methods for robust, transferable, decomposable, and spatially grounded concept representations. Moving beyond latent representations, we introduce external neurosymbolic reasoning procedures that construct explicit concept hierarchies and grounded logical rules around frozen vision-language models. We next study structure within model computation, identifying visual attention dilution in latent multi-agent reasoning and proposing an adaptive visual residual that preserves access to image evidence. Finally, we enforce structure through training. AƧttentive-CoT encourages delayed answer commitment and sustained visual access during multimodal chain-of-thought reasoning, while Chart-RVR decomposes reasoning into verifiable Structure, Evidence, and Derivation states optimized using programmatic rewards. Across these settings, the results show that interpretability and monitorability need not come at the expense of predictive performance. More broadly, this dissertation argues that trustworthy visual reasoning requires not merely exposing intermediate representations, but explicitly constraining what they contain, how they remain grounded in the input, and how their correctness can be independently checked.