Computer Vision & NLP
Hierarchical Co-Attention for Visual Question Answering
Jointly reasoning about visual attention ("where to look") and question attention ("which words to listen to") across word, phrase, and question hierarchies.

In visual question answering, a model must understand both the image and the language query. The difficulty is that the useful signal is often distributed: the image contains many irrelevant regions, while the question contains words that matter more than others.
This motivates co-attention: the system should learn not only where to look in the image, but also which words to prioritize in the question.

Why attention matters in VQA
A typical model may treat the question as a flat sequence and the image as a set of regional features. However, real reasoning is hierarchical:
- words combine into phrase-level meaning
- phrases compose into a question-level interpretation
- image regions compete for relevance
- the answer emerges from the joint alignment of both views
The hierarchical co-attention model addresses exactly this by aligning visual and textual representations at multiple levels.


Core idea
The architecture computes attention over two modalities simultaneously:
- visual attention decides which image regions matter most
- question attention decides which words or phrases are most relevant
This is particularly important in VQA because the same image may support multiple questions, and different query tokens may imply different regions of interest.
If the model only attends to one modality, it misses important context. The joint representation allows the question to guide the image interpretation and the image to guide the text interpretation.
Mathematical intuition
Let be the visual feature matrix and the question representation. The model builds an affinity matrix:
This matrix captures the compatibility between question features and visual features. From it, the model derives separate attention distributions over the image and the question.
The result is a richer representation where both views inform the final fused feature vector used for answer decoding.

Why hierarchical attention is powerful
A flat attention model may focus on the image as a whole and miss finer details. Hierarchical attention allows the reasoning process to operate at multiple scales:
- word-level alignment
- phrase-level fusion
- question-level summarization
- final image-text reasoning
This makes it especially effective for compositional questions such as "What is the person holding?" or "How many objects are visible to the left of the table?"

Practical takeaway
For multimodal systems, the key is not simply combining features. It is learning where the model should focus in the visual stream and what it should emphasize in the textual stream.
That is the core insight behind hierarchical co-attention: it allows the model to reason jointly across vision and language instead of treating them as independent inputs.
Closing note
Modern VQA systems have become far more sophisticated, but the core principle remains the same: the model must learn to align the right visual cues with the right textual cues. That alignment is exactly what co-attention aims to optimize.