Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Golden, Alicia, Hsia, Samuel, Sun, Fei, Acun, Bilge, Hosmer, Basil, Lee, Yejin, DeVito, Zachary, Johnson, Jeff, Wei, Gu-Yeon, Brooks, David, Wu, Carole-Jean |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024)
by: Golden, Alicia, et al.
Published: (2024)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023)
by: Hsia, Samuel, et al.
Published: (2023)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025)
by: Golden, Alicia, et al.
Published: (2025)
Beyond Efficiency: Scaling AI Sustainably
by: Wu, Carole-Jean, et al.
Published: (2024)
by: Wu, Carole-Jean, et al.
Published: (2024)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025)
by: Zhu, Jianian, et al.
Published: (2025)
STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data
by: Wang, Daoce, et al.
Published: (2025)
by: Wang, Daoce, et al.
Published: (2025)
PPVF: An Efficient Privacy-Preserving Online Video Fetching Framework with Correlated Differential Privacy
by: Zhang, Xianzhi, et al.
Published: (2024)
by: Zhang, Xianzhi, et al.
Published: (2024)
EdgeVision: Towards Collaborative Video Analytics on Distributed Edges for Performance Maximization
by: Gao, Guanyu, et al.
Published: (2022)
by: Gao, Guanyu, et al.
Published: (2022)
Steganography -- coding and intercepting the information from encoded pictures in the absence of any initial information
by: Kwiatkowska, Monika, et al.
Published: (2014)
by: Kwiatkowska, Monika, et al.
Published: (2014)
$2B$ or Not $2B$: A Tale of Three Algorithms for Streaming: Covariance Estimation after Welford and Chan-Golub-LeVeque
by: Reichel, Felix
Published: (2026)
by: Reichel, Felix
Published: (2026)
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Stimpack: An Adaptive Rendering Optimization System for Scalable Cloud Gaming
by: Heo, Jin, et al.
Published: (2024)
by: Heo, Jin, et al.
Published: (2024)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
by: Li, Minghao, et al.
Published: (2026)
by: Li, Minghao, et al.
Published: (2026)
Towards Real-Time Neural Volumetric Rendering on Mobile Devices: A Measurement Study
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
AFLL: Real-time Load Stabilization for MMO Game Servers Based on Circular Causality Learning
by: Kang, Shinsuk, et al.
Published: (2026)
by: Kang, Shinsuk, et al.
Published: (2026)
HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy Preservation in Multimodal Sentiment Analysis
by: Wu, Zhuojia, et al.
Published: (2024)
by: Wu, Zhuojia, et al.
Published: (2024)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Health+: Empowering Individuals via Unifying Health Data
by: Maiyya, Sujaya, et al.
Published: (2026)
by: Maiyya, Sujaya, et al.
Published: (2026)
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
by: Xiong, Zechang, et al.
Published: (2026)
by: Xiong, Zechang, et al.
Published: (2026)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)
by: Kokolis, Apostolos, et al.
Published: (2024)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
by: Zhang, Meishan, et al.
Published: (2024)
by: Zhang, Meishan, et al.
Published: (2024)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
by: Hu, Jiayun, et al.
Published: (2025)
by: Hu, Jiayun, et al.
Published: (2025)
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models
by: Chen, Haokun, et al.
Published: (2024)
by: Chen, Haokun, et al.
Published: (2024)
Producer vs. Rapper: Who Dominates the Hip Hop Sound? A Case Study
by: Ziemer, Tim, et al.
Published: (2024)
by: Ziemer, Tim, et al.
Published: (2024)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
by: Li, Shuyu, et al.
Published: (2025)
by: Li, Shuyu, et al.
Published: (2025)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
by: Xie, Tianyidan, et al.
Published: (2026)
by: Xie, Tianyidan, et al.
Published: (2026)
Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering
by: Fei, Ben, et al.
Published: (2024)
by: Fei, Ben, et al.
Published: (2024)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
by: Yang, Quanwei, et al.
Published: (2025)
by: Yang, Quanwei, et al.
Published: (2025)
Toward Accessible and Safe Live Streaming Using Distributed Content Filtering with MoQ
by: Freeman, Andrew C.
Published: (2025)
by: Freeman, Andrew C.
Published: (2025)
SFedCA: Credit Assignment-Based Active Client Selection Strategy for Spiking Federated Learning
by: Zhan, Qiugang, et al.
Published: (2024)
by: Zhan, Qiugang, et al.
Published: (2024)
LongCat-Flash-Omni Technical Report
by: Meituan LongCat Team, et al.
Published: (2025)
by: Meituan LongCat Team, et al.
Published: (2025)
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
by: Du, Kuntai, et al.
Published: (2023)
by: Du, Kuntai, et al.
Published: (2023)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
by: Li, Yaoru, et al.
Published: (2025)
by: Li, Yaoru, et al.
Published: (2025)
ConCLVD: Controllable Chinese Landscape Video Generation via Diffusion Model
by: Liu, Dingming, et al.
Published: (2024)
by: Liu, Dingming, et al.
Published: (2024)
Generative Semantic Communication: Diffusion Models Beyond Bit Recovery
by: Grassucci, Eleonora, et al.
Published: (2023)
by: Grassucci, Eleonora, et al.
Published: (2023)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
by: Wei, Tianxin, et al.
Published: (2024)
by: Wei, Tianxin, et al.
Published: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
by: Zou, Heqing, et al.
Published: (2024)
by: Zou, Heqing, et al.
Published: (2024)
Similar Items
-
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024) -
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023) -
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025) -
Beyond Efficiency: Scaling AI Sustainably
by: Wu, Carole-Jean, et al.
Published: (2024) -
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025)