Multimodal Representation
Alignment, fusion and transferable representations across heterogeneous modalities.
Alignment, fusion and transferable representations across heterogeneous modalities.
Robust perception and state estimation with complementary visual and geometric sensors.
Structured multimodal learning for 3D scenes and scientific observations.
A multi-sensor perception system integrating RGB cameras and LiDAR for online localization, state estimation and closed-loop evaluation.
Exploring representations that connect dense 3D seismic volumes with sparse well-log observations.
Notes on decoupled alignment in multimodal representation learning and what it changes about shared representations.
A short reading note on modality-generalizable representations, shared structure and domain shift.
A compact derivation-oriented note on 3D Gaussian primitives, covariance parameterization and projection.
Practical observations from result-level fusion experiments in a multi-sensor localization pipeline.