LinMU: Multimodal Understanding Made Linear
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hongjie, Jha, Niraj K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers
von: Wang, Hongjie, et al.
Veröffentlicht: (2023)
von: Wang, Hongjie, et al.
Veröffentlicht: (2023)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
Understanding Audiovisual Deepfake Detection: Techniques, Challenges, Human Factors and Perceptual Insights
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2024)
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2024)
Semantic-Aware Adaptive Video Streaming Using Latent Diffusion Models for Wireless Networks
von: Yan, Zijiang, et al.
Veröffentlicht: (2025)
von: Yan, Zijiang, et al.
Veröffentlicht: (2025)
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
RAPNet: A Receptive-Field Adaptive Convolutional Neural Network for Pansharpening
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Mamba-360: Survey of State Space Models as Transformer Alternative for Long Sequence Modelling: Methods, Applications, and Challenges
von: Patro, Badri Narayana, et al.
Veröffentlicht: (2024)
von: Patro, Badri Narayana, et al.
Veröffentlicht: (2024)
Food Portion Estimation via 3D Object Scaling
von: Vinod, Gautham, et al.
Veröffentlicht: (2024)
von: Vinod, Gautham, et al.
Veröffentlicht: (2024)
VEMOCLAP: A video emotion classification web application
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
Health AI Developer Foundations
von: Kiraly, Atilla P., et al.
Veröffentlicht: (2024)
von: Kiraly, Atilla P., et al.
Veröffentlicht: (2024)
Improving Multi-label Recognition using Class Co-Occurrence Probabilities
von: Rawlekar, Samyak, et al.
Veröffentlicht: (2024)
von: Rawlekar, Samyak, et al.
Veröffentlicht: (2024)
FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait
von: Ki, Taekyung, et al.
Veröffentlicht: (2024)
von: Ki, Taekyung, et al.
Veröffentlicht: (2024)
MIND: A Noise-Adaptive Denoising Framework for Medical Images Integrating Multi-Scale Transformer
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Scaling Up Single Image Dehazing Algorithm by Cross-Data Vision Alignment for Richer Representation Learning and Beyond
von: Shi, Yukai, et al.
Veröffentlicht: (2024)
von: Shi, Yukai, et al.
Veröffentlicht: (2024)
Attention-Driven Training-Free Efficiency Enhancement of Diffusion Models
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
Movie Trailer Genre Classification Using Multimodal Pretrained Features
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
von: Henschel, Roberto, et al.
Veröffentlicht: (2024)
von: Henschel, Roberto, et al.
Veröffentlicht: (2024)
SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation
von: Wang, Qizhou, et al.
Veröffentlicht: (2026)
von: Wang, Qizhou, et al.
Veröffentlicht: (2026)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026)
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026)
Spatial Visibility and Temporal Dynamics: Revolutionizing Field of View Prediction in Adaptive Point Cloud Video Streaming
von: Li, Chen, et al.
Veröffentlicht: (2024)
von: Li, Chen, et al.
Veröffentlicht: (2024)
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Food Portion Estimation: From Pixels to Calories
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
Towards Real-world Video Face Restoration: A New Benchmark
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
Machine Perception-Driven Image Compression: A Layered Generative Approach
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2023)
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2023)
Change Detection Between Optical Remote Sensing Imagery and Map Data via Segment Anything Model (SAM)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2024)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2024)
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
PixelBoost: Leveraging Brownian Motion for Realistic-Image Super-Resolution
von: Mishra, Aradhana, et al.
Veröffentlicht: (2025)
von: Mishra, Aradhana, et al.
Veröffentlicht: (2025)
RAISE: Realness Assessment for Image Synthesis and Evaluation
von: Mukherjee, Aniruddha, et al.
Veröffentlicht: (2025)
von: Mukherjee, Aniruddha, et al.
Veröffentlicht: (2025)
From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision
von: Adhikarla, Eashan, et al.
Veröffentlicht: (2024)
von: Adhikarla, Eashan, et al.
Veröffentlicht: (2024)
Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder Estimation
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
HPC: Hierarchical Progressive Coding Framework for Volumetric Video
von: Zheng, Zihan, et al.
Veröffentlicht: (2024)
von: Zheng, Zihan, et al.
Veröffentlicht: (2024)
FineVQ: Fine-Grained User Generated Content Video Quality Assessment
von: Duan, Huiyu, et al.
Veröffentlicht: (2024)
von: Duan, Huiyu, et al.
Veröffentlicht: (2024)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
von: Nasir, Tayyab, et al.
Veröffentlicht: (2026)
von: Nasir, Tayyab, et al.
Veröffentlicht: (2026)
SCENE: Semantic-aware Codec Enhancement with Neural Embeddings
von: Lin, Han-Yu, et al.
Veröffentlicht: (2026)
von: Lin, Han-Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers
von: Wang, Hongjie, et al.
Veröffentlicht: (2023) -
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
von: Wang, Hongjie, et al.
Veröffentlicht: (2024) -
Understanding Audiovisual Deepfake Detection: Techniques, Challenges, Human Factors and Perceptual Insights
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2024) -
Semantic-Aware Adaptive Video Streaming Using Latent Diffusion Models for Wireless Networks
von: Yan, Zijiang, et al.
Veröffentlicht: (2025) -
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)