Prof. Paul Liang lecturing on multimodal AI foundations and core challenges at the 2026 Jeju Summer Camp

Multimodal AI: Foundations, Core Challenges, and Future Directions (I)

SCROLL

Multimodal AI: Foundations, Core Challenges, and Future Directions (I)



Multimodal AI: Foundations, Core Challenges, and Future Directions (I)

  • Speaker : Prof. Paul Liang
  • Date : July 6, 2026
  • Affiliation : Assistant Professor, MIT Media Lab & MIT EECS, Massachusetts Institute of Technology, USA
  • Category : Special Lecture (2026 Jeju Summer Camp)

Multimodal AI: Foundations, Core Challenges, and Future Directions (I)

Abstract

This special lecture at the 2026 Jeju Summer Camp gives a broad overview of multimodal AI, the science of learning from heterogeneous and interconnected data such as language, vision, audio, touch, and physiological signals. The framing is deliberate. Multimodality is not simply a matter of accepting more input types, so the lecture asks what makes a modality in the first place, why modalities are heterogeneous yet connected, and how they interact.

Three kinds of interaction organize that answer: redundancy, where both modalities carry the same information, uniqueness, where only one does, and synergy, where meaning emerges that neither reveals alone. Sarcasm is the example for the last. On that footing the lecture lays out six core technical challenges, then traces how today’s large multimodal models address them, from multimodal transformers and contrastive pre-training to conditioning frozen language models for vision-language reasoning and generation.

The closing part looks past text and images to research frontiers: systems that sense touch and smell, and agents that grow their own memory, reasoning, and capabilities over time.

Presentation Overview

This presentation covers the following key topics:

  • Why multimodal AI matters: multimedia and content creation, physical sensing for manufacturing and robotics, holistic health across physical, social, and emotional wellbeing, and digital agents for the web and OS-level automation
  • Defining the field: what counts as a modality, raw against abstract and near against far from the sensor, and a research definition of multimodal as data that is heterogeneous, connected, and interacting
  • Modality interactions: redundancy as information shared by both modalities, uniqueness as information carried by only one, and synergy as emergent meaning such as sarcasm
  • The six core challenges: representation (fusion, coordination, fission), alignment (discrete, continuous, aligned representation), reasoning (multi-step inference with external knowledge), generation (summarization, translation, creation), transference (transfer, co-learning, model induction), and quantification (heterogeneity, interactions, learning dynamics)
  • Large multimodal models: multimodal transformers and cross-modal attention, ViLT, and CLIP-style contrastive language and image pre-training with the InfoNCE objective for coordinated representation spaces
  • Adapting language models into multimodal ones: prefix tuning and conditioning on visual input in Frozen, Flamingo, MiniGPT-4’s two-stage alignment and instruction tuning, and LLaMA-Adapter
  • Data at scale: the path from YFCC-100M and LAION to DataComp-12B, the turn toward quality filtering, and the more fragmented landscape of multimodal instruction-tuning datasets in general and clinical domains
  • From text to multimodal generation: diffusion models conditioned on CLIP latents for text-to-image generation, and grounding language models for interleaved image retrieval and image generation
  • New modalities, touch: low-cost tactile sensing gloves, the OpenTouch dataset of full-hand touch with egocentric vision and 3D hand pose, vision and touch retrieval, and vision-tactile-action models fusing language, vision, and tactile streams sampled at very different rates for fast reactive control
  • New modalities, smell: SmellNet as a large-scale real-world smell recognition dataset, cross-modal alignment using molecular databases as a stronger training-time modality, allergen detection, and AromaGen for interactive olfactory generation
  • Self-evolving multimodal AI: MEM1’s internal-state memory consolidation for near-constant-memory long-horizon agents, Propose-Solve-Verify self-play through formal verification, and CORAL’s self-evolving multi-agent systems with shared memory and specialization
  • Takeaways: multimodal problems are defined by heterogeneity, connection, and interaction, the six challenges map the search for the right method for a given application, and the frontier is moving toward new senses, multimodal agents, and AI that improves itself recursively

© 2023 KAIST Mobility Group.
All rights reserved.

Sitemap

T : +82 42 350 1252
F :+82 42 350 1250

E : kaist.mobility@gmail.com