Index of papers
131 reviewed papers · 31 tags
-
Chameleon: Mixed-Modal Early-Fusion Foundation Models
-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
-
Mixtral of Experts
-
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
-
3D Gaussian Splatting for Real-Time Radiance Field Rendering
-
Adding Conditional Control to Text-to-Image Diffusion Models
-
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
-
Consistency Models
-
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
-
Efficient Memory Management for Large Language Model Serving with PagedAttention
-
Fast Inference from Transformers via Speculative Decoding
-
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
-
Generative Inverse Design of Metamaterials with Functional Responses by Interpretable Learning
-
Let's Verify Step by Step
-
LLaMA: Open and Efficient Foundation Language Models
-
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
-
Mastering Diverse Domains through World Models
-
NExT-GPT: Any-to-Any Multimodal LLM
-
QLoRA: Efficient Finetuning of Quantized LLMs
-
RWKV: Reinventing RNNs for the Transformer Era
-
Segment Anything
-
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
-
Toolformer: Language Models Can Teach Themselves to Use Tools
-
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
-
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
-
Classifier-Free Diffusion Guidance
-
Constitutional AI: Harmlessness from AI Feedback
-
DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps
-
Elucidating the Design Space of Diffusion-Based Generative Models
-
Flamingo: a Visual Language Model for Few-Shot Learning
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
-
Flow Matching for Generative Modeling
-
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
-
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
-
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
-
High-Resolution Image Synthesis with Latent Diffusion Models
-
Large Language Models are Zero-Shot Reasoners
-
ReAct: Synergizing Reasoning and Acting in Language Models
-
Robust Speech Recognition via Large-Scale Weak Supervision
-
Scalable Diffusion Models with Transformers
-
Training Compute-Optimal Large Language Models
-
Training Language Models to Follow Instructions with Human Feedback
-
Video Diffusion Models
-
Decision Transformer: Reinforcement Learning via Sequence Modeling
-
Diffusion Models Beat GANs on Image Synthesis
-
Efficiently Modeling Long Sequences with Structured State Spaces
-
Emerging Properties in Self-Supervised Vision Transformers
-
Evaluating Large Language Models Trained on Code
-
Improved Denoising Diffusion Probabilistic Models
-
Learning Transferable Visual Models from Natural Language Supervision
-
LoRA: Low-Rank Adaptation of Large Language Models
-
Masked Autoencoders Are Scalable Vision Learners
-
Mastering Atari with Discrete World Models
-
RoFormer: Enhanced Transformer with Rotary Position Embedding
-
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
-
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
-
Zero-Shot Text-to-Image Generation
-
A Simple Framework for Contrastive Learning of Visual Representations
-
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
-
Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning
-
Denoising Diffusion Implicit Models
-
Denoising Diffusion Probabilistic Models
-
End-to-End Object Detection with Transformers
-
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
-
Language Models are Few-Shot Learners
-
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
-
Momentum Contrast for Unsupervised Visual Representation Learning
-
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
-
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
-
Scaling Laws for Neural Language Models
-
Score-Based Generative Modeling through Stochastic Differential Equations
-
The Curious Case of Neural Text Degeneration
-
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
-
A Style-Based Generator Architecture for Generative Adversarial Networks
-
Decoupled Weight Decay Regularization
-
Dream to Control: Learning Behaviors by Latent Imagination
-
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
-
Language Models are Unsupervised Multitask Learners
-
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
-
A General Reinforcement Learning Algorithm that Masters Chess, Shogi, and Go through Self-Play
-
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
-
Deep Contextualized Word Representations
-
Graph Attention Networks
-
Group Normalization
-
Improving Language Understanding by Generative Pre-Training
-
Recurrent World Models Facilitate Policy Evolution
-
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
-
Attention Is All You Need
-
Deep Reinforcement Learning from Human Preferences
-
Densely Connected Convolutional Networks
-
Density Estimation using Real NVP
-
Inductive Representation Learning on Large Graphs
-
Instance Normalization: The Missing Ingredient for Fast Stylization
-
Mask R-CNN
-
Proximal Policy Optimization Algorithms
-
Semi-Supervised Classification with Graph Convolutional Networks
-
Asynchronous Methods for Deep Reinforcement Learning
-
Deep Residual Learning for Image Recognition
-
Layer Normalization
-
Mastering the Game of Go with Deep Neural Networks and Tree Search
-
Neural Machine Translation of Rare Words with Subword Units
-
WaveNet: A Generative Model for Raw Audio
-
You Only Look Once: Unified, Real-Time Object Detection
-
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
-
Continuous control with deep reinforcement learning
-
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
-
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
-
Distilling the Knowledge in a Neural Network
-
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
-
Going Deeper with Convolutions
-
Human-Level Control through Deep Reinforcement Learning
-
Trust Region Policy Optimization
-
U-Net: Convolutional Networks for Biomedical Image Segmentation
-
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
-
Adam: A Method for Stochastic Optimization
-
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
-
Generative Adversarial Networks
-
GloVe: Global Vectors for Word Representation
-
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
-
Neural Machine Translation by Jointly Learning to Align and Translate
-
Sequence to Sequence Learning with Neural Networks
-
Very Deep Convolutional Networks for Large-Scale Image Recognition
-
Auto-Encoding Variational Bayes
-
Efficient Estimation of Word Representations in Vector Space
-
Representation Learning: A Review and New Perspectives
-
ImageNet Classification with Deep Convolutional Neural Networks
-
A Fast Learning Algorithm for Deep Belief Nets
-
Gradient-based learning applied to document recognition
-
Long Short-Term Memory
-
Learning Representations by Back-propagating Errors
-
The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain