LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Hanyu, Lee, Gim Hee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging Scenes
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Segment Any Events with Language
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026)
by: Zhou, Hanyu, et al.
Published: (2026)
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational Frequencies
by: Lu, Dongyue, et al.
Published: (2024)
by: Lu, Dongyue, et al.
Published: (2024)
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting
by: Zhou, Haoran, et al.
Published: (2026)
by: Zhou, Haoran, et al.
Published: (2026)
DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting
by: Lee, Seungjun, et al.
Published: (2025)
by: Lee, Seungjun, et al.
Published: (2025)
Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
Event-Driven Dynamic Scene Depth Completion
by: Yan, Zhiqiang, et al.
Published: (2025)
by: Yan, Zhiqiang, et al.
Published: (2025)
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
by: Li, Geng, et al.
Published: (2025)
by: Li, Geng, et al.
Published: (2025)
Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE
by: Bao, Wei, et al.
Published: (2026)
by: Bao, Wei, et al.
Published: (2026)
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
Unified Geometry and Color Compression Framework for Point Clouds via Generative Diffusion Priors
by: Huang, Tianxin, et al.
Published: (2025)
by: Huang, Tianxin, et al.
Published: (2025)
Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM
by: Wang, Yunsong, et al.
Published: (2026)
by: Wang, Yunsong, et al.
Published: (2026)
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
by: Cheng, Wencan, et al.
Published: (2026)
by: Cheng, Wencan, et al.
Published: (2026)
DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian Consensus
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
by: Zhang, Can, et al.
Published: (2026)
by: Zhang, Can, et al.
Published: (2026)
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
by: Deng, Jiajun, et al.
Published: (2025)
by: Deng, Jiajun, et al.
Published: (2025)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Instance-Level Moving Object Segmentation from a Single Image with Events
by: Wan, Zhexiong, et al.
Published: (2025)
by: Wan, Zhexiong, et al.
Published: (2025)
Deblur e-NeRF: NeRF from Motion-Blurred Events under High-speed or Low-light Conditions
by: Low, Weng Fei, et al.
Published: (2024)
by: Low, Weng Fei, et al.
Published: (2024)
NEC-Diff: Noise-Robust Event-RAW Complementary Diffusion for Seeing Motion in Extreme Darkness
by: Liu, Haoyue, et al.
Published: (2026)
by: Liu, Haoyue, et al.
Published: (2026)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
Bring Event into RGB and LiDAR: Hierarchical Visual-Motion Fusion for Scene Flow
by: Zhou, Hanyu, et al.
Published: (2024)
by: Zhou, Hanyu, et al.
Published: (2024)
Syn-to-Real Unsupervised Domain Adaptation for Indoor 3D Object Detection
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic Fields
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
ChatSplat: 3D Conversational Gaussian Splatting
by: Chen, Hanlin, et al.
Published: (2024)
by: Chen, Hanlin, et al.
Published: (2024)
DiHuR: Diffusion-Guided Generalizable Human Reconstruction
by: Chen, Jinnan, et al.
Published: (2024)
by: Chen, Jinnan, et al.
Published: (2024)
Similar Items
-
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025) -
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
by: Zhou, Hanyu, et al.
Published: (2025) -
Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging Scenes
by: Wang, Haonan, et al.
Published: (2025) -
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
by: Zhou, Hanyu, et al.
Published: (2025) -
Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow
by: Zhou, Hanyu, et al.
Published: (2025)