Characterizing and Efficiently Accelerating Multimodal Generation Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yejin, Sun, Anna, Hosmer, Basil, Acun, Bilge, Balioglu, Can, Wang, Changhan, Hernandez, Charles David, Puhrsch, Christian, Haziza, Daniel, Guessous, Driss, Massa, Francisco, Kahn, Jacob, Wan, Jeffrey, Reizenstein, Jeremy, Zhai, Jiaqi, Isaacson, Joe, Schlosser, Joel, Pino, Juan, Sadagopan, Kaushik Ram, Shamis, Leonid, Ma, Linjian, Hwang, Min-Jae, Chen, Mingda, Elhoushi, Mostafa, Rodriguez, Pedro, Pasunuru, Ram, Yih, Scott, Popuri, Sravya, Liu, Xing, Wu, Carole-Jean |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024)
by: Elhoushi, Mostafa, et al.
Published: (2024)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Diagonal operators, $q$-Whittaker functions and rook theory
by: Ram, Samrith, et al.
Published: (2023)
by: Ram, Samrith, et al.
Published: (2023)
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024)
by: Golden, Alicia, et al.
Published: (2024)
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
by: Golden, Alicia, et al.
Published: (2023)
by: Golden, Alicia, et al.
Published: (2023)
NeurIPS 2023 LLM Efficiency Fine-tuning Competition
by: Saroufim, Mark, et al.
Published: (2025)
by: Saroufim, Mark, et al.
Published: (2025)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
Designing a Future-Proof Cryptography Strategy for Quantum Computing
by: Sreekanth Pasunuru
Published: (2024)
by: Sreekanth Pasunuru
Published: (2024)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
by: Wang, Irene, et al.
Published: (2025)
by: Wang, Irene, et al.
Published: (2025)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
by: Dong, Harry, et al.
Published: (2025)
by: Dong, Harry, et al.
Published: (2025)
Hosmer, Susan C Feb. 18, 1915 [Postcard]
by: Hosmer, Susan C
Published: (1915)
by: Hosmer, Susan C
Published: (1915)
TorchAO: PyTorch-Native Training-to-Serving Model Optimization
by: Or, Andrew, et al.
Published: (2025)
by: Or, Andrew, et al.
Published: (2025)
Beyond Efficiency: Scaling AI Sustainably
by: Wu, Carole-Jean, et al.
Published: (2024)
by: Wu, Carole-Jean, et al.
Published: (2024)
Investigating Decoder-only Large Language Models for Speech-to-text Translation
by: Huang, Chao-Wei, et al.
Published: (2024)
by: Huang, Chao-Wei, et al.
Published: (2024)
Support theories for non-Noetherian tensor triangulated categories
by: Zou, Changhan
Published: (2023)
by: Zou, Changhan
Published: (2023)
Convexity in tensor triangular geometry
by: Zou, Changhan
Published: (2025)
by: Zou, Changhan
Published: (2025)
Terminal Steiner tree problem : Complexity and Algorithms
by: S, Jyothish, et al.
Published: (2026)
by: S, Jyothish, et al.
Published: (2026)
Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering
by: Chen, Mingda, et al.
Published: (2023)
by: Chen, Mingda, et al.
Published: (2023)
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
by: Guo, Han, et al.
Published: (2026)
by: Guo, Han, et al.
Published: (2026)
Enhancing Gluten‐Free Cake Quality With Germinated and Ungerminated Mung Bean Flours: A Comparative Analysis
by: Sultan Acun
Published: (2026)
by: Sultan Acun
Published: (2026)
Unlocking the Potential of Renewable Energy Through Curtailment Prediction
by: Acun, Bilge, et al.
Published: (2024)
by: Acun, Bilge, et al.
Published: (2024)
Secure Domination in Bisplit graphs -- A Structural and algorithmic study
by: D, Swathi, et al.
Published: (2025)
by: D, Swathi, et al.
Published: (2025)
Set Block Decoding is a Language Model Inference Accelerator
by: Gat, Itai, et al.
Published: (2025)
by: Gat, Itai, et al.
Published: (2025)
Machine learning methods for finite population parameter estimation in survey sampling
by: Dagdoug, Mehdi, et al.
Published: (2026)
by: Dagdoug, Mehdi, et al.
Published: (2026)
Area Law for the entanglement entropy of free fermions in nonrandom ergodic field
by: Pastur, Leonid, et al.
Published: (2025)
by: Pastur, Leonid, et al.
Published: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
by: Elhoushi, Mostafa, et al.
Published: (2025)
by: Elhoushi, Mostafa, et al.
Published: (2025)
LONER: LiDAR Only Neural Representations for Real-Time SLAM
by: Isaacson, Seth, et al.
Published: (2023)
by: Isaacson, Seth, et al.
Published: (2023)
Memory Layers at Scale
by: Berges, Vincent-Pierre, et al.
Published: (2024)
by: Berges, Vincent-Pierre, et al.
Published: (2024)
Conjugation, loop and closure invariants of the iterated-integrals signature
by: Diehl, Joscha, et al.
Published: (2024)
by: Diehl, Joscha, et al.
Published: (2024)
Disruptive Ecologies: Design with Nonhuman Intelligences
by: Roberto Bottazzi, et al.
Published: (2024)
by: Roberto Bottazzi, et al.
Published: (2024)
US Home Improvement Industry Benchmarks & Trends Report
by: KL, Ram
Published: (2026)
by: KL, Ram
Published: (2026)
Impact of high stocking density on growth parameters, amino acid profile and fatty acid profile of different fish species: a review
by: Ram, sha
Published: (2020)
by: Ram, sha
Published: (2020)
Spectral Operator Framework for Riemann Hypothesis
by: Ram, Pranad
Published: (2025)
by: Ram, Pranad
Published: (2025)
Analysis, Identification and Prediction of Parkinson Disease Sub-Types and Progression through Machine Learning
by: Ram, Ashwin
Published: (2023)
by: Ram, Ashwin
Published: (2023)
Subspace Profiles over Finite Fields and $q$-Whittaker Expansions of Symmetric Functions
by: Ram, Samrith
Published: (2023)
by: Ram, Samrith
Published: (2023)
Simple operators and $q$-Whittaker coefficients of power sum symmetric functions
by: Ram, Samrith
Published: (2024)
by: Ram, Samrith
Published: (2024)
Investigating Plausibility of Biologically Inspired Bayesian Learning in ANNs
by: Zaveri, Ram
Published: (2024)
by: Zaveri, Ram
Published: (2024)
Similar Items
-
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024) -
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
by: Peng, Yifan, et al.
Published: (2024) -
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
by: Peng, Yifan, et al.
Published: (2024) -
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
by: Elhoushi, Mostafa, et al.
Published: (2024) -
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)