Prof. Paul Liang lecturing on large multimodal models, new senses, and self-evolving agents at the 2026 Jeju Summer Camp

Multimodal AI: Foundations, Core Challenges, and Future Directions (II)

SCROLL

Multimodal AI: Foundations, Core Challenges, and Future Directions (II)



Multimodal AI: Foundations, Core Challenges, and Future Directions (II)

  • Speaker : Prof. Paul Liang
  • Date : July 6, 2026
  • Affiliation : Assistant Professor, MIT Media Lab & MIT EECS, Massachusetts Institute of Technology, USA
  • Category : Special Lecture (2026 Jeju Summer Camp)

Multimodal AI: Foundations, Core Challenges, and Future Directions (II)

Abstract

This is the second half of Prof. Paul Liang’s two-part lecture at the 2026 Jeju Summer Camp. Part I set up the foundations, defining what a modality is and how modalities interact through redundancy, uniqueness, and synergy, then mapped the field onto six core challenges. Read it first at Multimodal AI: Foundations, Core Challenges, and Future Directions (I).

This session carries that map into practice. It follows how today’s large multimodal models take on those challenges, from multimodal transformers and cross-modal attention through contrastive language and image pre-training, then on to the methods that turn a frozen language model into a multimodal one. Data is treated as a first-class concern, from web-scale image and text corpora to the more fragmented world of multimodal instruction tuning.

The closing stretch moves past text and images. Tactile sensing and smell recognition bring new senses into reach, and self-evolving agents consolidate their own memory, verify their own work, and specialize as a group. The takeaway is that heterogeneity, connection, and interaction define these problems, and that the frontier now runs through new senses, multimodal agents, and systems that improve themselves.

Presentation Overview

The lecture ran across two sessions and the recordings share one description, so the split below follows the running order of the source: Part I covers the foundations and the six challenges, and this session picks up from the models onward.

  • Large multimodal models: multimodal transformers and cross-modal attention, ViLT, and CLIP-style contrastive language and image pre-training with the InfoNCE objective for coordinated representation spaces
  • Adapting language models into multimodal ones: prefix tuning and conditioning on visual input in Frozen, Flamingo, MiniGPT-4’s two-stage alignment and instruction tuning, and LLaMA-Adapter
  • Data at scale: the path from YFCC-100M and LAION to DataComp-12B, the turn toward quality filtering, and the more fragmented landscape of multimodal instruction-tuning datasets in general and clinical domains
  • From text to multimodal generation: diffusion models conditioned on CLIP latents for text-to-image generation, and grounding language models for interleaved image retrieval and image generation
  • New modalities, touch: low-cost tactile sensing gloves, the OpenTouch dataset of full-hand touch with egocentric vision and 3D hand pose, vision and touch retrieval, and vision-tactile-action models fusing language, vision, and tactile streams sampled at very different rates for fast reactive control
  • New modalities, smell: SmellNet as a large-scale real-world smell recognition dataset, cross-modal alignment using molecular databases as a stronger training-time modality, allergen detection, and AromaGen for interactive olfactory generation
  • Self-evolving multimodal AI: MEM1’s internal-state memory consolidation for near-constant-memory long-horizon agents, Propose-Solve-Verify self-play through formal verification, and CORAL’s self-evolving multi-agent systems with shared memory and specialization
  • Takeaways: multimodal problems are defined by heterogeneity, connection, and interaction, the six challenges map the search for the right method for a given application, and the frontier is moving toward new senses, multimodal agents, and AI that improves itself recursively

© 2023 KAIST Mobility Group.
All rights reserved.

Sitemap

T : +82 42 350 1252
F :+82 42 350 1250

E : kaist.mobility@gmail.com