Papers tagged “efficiency”
18 papers · All papers →
-
Mixtral of Experts
-
3D Gaussian Splatting for Real-Time Radiance Field Rendering
-
Consistency Models
-
Efficient Memory Management for Large Language Model Serving with PagedAttention
-
Fast Inference from Transformers via Speculative Decoding
-
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
-
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
-
QLoRA: Efficient Finetuning of Quantized LLMs
-
RWKV: Reinventing RNNs for the Transformer Era
-
DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
-
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
-
Efficiently Modeling Long Sequences with Structured State Spaces
-
LoRA: Low-Rank Adaptation of Large Language Models
-
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
-
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
-
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
-
Distilling the Knowledge in a Neural Network