Multimodal Learning

2 min
  1. Which unit must stay whole when splitting multimodal data?
  2. What does the control where image–text pairs are randomly shuffled verify?
  3. Why remove one modality during evaluation?
  4. What behavior must be defined when a sensor is missing in production?
  5. Does a higher global performance prove that all modalities are useful?

Answers. 1. The entity that connects the views — scene, person, session or document — so that another view of the same case does not cross partitions. 2. That the gain depends on a real alignment rather than on unimodal shortcuts or class frequency. 3. An ablation measures the contribution of each source and reveals a decorative modality. 4. An explicit degraded mode, abstention or tested fallback. 5. No: ablations, contradictory cases, missing inputs and comparison to the best unimodal model at the same budget are required.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Modern Deep Dives” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one