SimMMDG
Question
What does a representation need to capture if we expect it to transfer across both domains and modalities?
Notes
The useful distinction for me is between cross-modal commonality and sensor-dependent evidence. A representation that only compresses common information may become stable, but it can also become weak for tasks that rely on modality-specific cues.
Connection to my work
This is relevant to sensor fusion systems where one modality may degrade or disappear.