Note · SEP 12, 2026

DecAlign — Reading Notes

Notes on decoupled alignment in multimodal representation learning and what it changes about shared representations.

Multimodal LearningAlignmentRepresentation

DecAlign — Reading Notes

Why I read this paper

I am interested in a recurring question in multimodal learning: what should be shared across modalities, and what should remain modality-specific?

Problem

Heterogeneous sensors observe complementary properties of the same physical world, so forcing every feature into a fully shared space may discard useful modality-specific information.

What I want to remember

My thoughts

For RGB–LiDAR perception, image appearance and point-cloud geometry are correlated but not interchangeable. A useful representation should support cross-modal correspondence without erasing complementary structure.