Papers tagged “multimodal”
7 papers · All papers →
-
Chameleon: Mixed-Modal Early-Fusion Foundation Models
-
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
-
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
-
NExT-GPT: Any-to-Any Multimodal LLM
-
Flamingo: a Visual Language Model for Few-Shot Learning
-
Learning Transferable Visual Models from Natural Language Supervision
-
Zero-Shot Text-to-Image Generation