LLM4Brain: Training a Large Language Model for Brain Video Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Ruizhe, Sun, Lichao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MindCross: Fast New Subject Adaptation with Limited Data for Cross-subject Video Reconstruction from Brain Signals
por: Liu, Xuan-Hao, et al.
Publicado: (2025)
por: Liu, Xuan-Hao, et al.
Publicado: (2025)
Deep Neural Encoder-Decoder Model to Relate fMRI Brain Activity with Naturalistic Stimuli
por: David, Florian, et al.
Publicado: (2025)
por: David, Florian, et al.
Publicado: (2025)
Adaptive Modality Balanced Online Knowledge Distillation for Brain-Eye-Computer based Dim Object Detection
por: Li, Zixing, et al.
Publicado: (2024)
por: Li, Zixing, et al.
Publicado: (2024)
ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding
por: Liu, Minxu, et al.
Publicado: (2025)
por: Liu, Minxu, et al.
Publicado: (2025)
Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
por: Zhang, Guowei, et al.
Publicado: (2025)
por: Zhang, Guowei, et al.
Publicado: (2025)
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
por: Zhou, Tian-Yi, et al.
Publicado: (2026)
por: Zhou, Tian-Yi, et al.
Publicado: (2026)
Reframe Anything: LLM Agent for Open World Video Reframing
por: Cao, Jiawang, et al.
Publicado: (2024)
por: Cao, Jiawang, et al.
Publicado: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
por: Sun, Boyuan, et al.
Publicado: (2026)
por: Sun, Boyuan, et al.
Publicado: (2026)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
por: Fei, Hao, et al.
Publicado: (2024)
por: Fei, Hao, et al.
Publicado: (2024)
ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
por: Xia, Ding, et al.
Publicado: (2025)
por: Xia, Ding, et al.
Publicado: (2025)
Acoustic Field Video for Multimodal Scene Understanding
por: Kim, Daehwa, et al.
Publicado: (2026)
por: Kim, Daehwa, et al.
Publicado: (2026)
When, Where, and What? A Novel Benchmark for Accident Anticipation and Localization with Large Language Models
por: Liao, Haicheng, et al.
Publicado: (2024)
por: Liao, Haicheng, et al.
Publicado: (2024)
Rehabilitation Exercise Quality Assessment and Feedback Generation Using Large Language Models with Prompt Engineering
por: Tang, Jessica, et al.
Publicado: (2025)
por: Tang, Jessica, et al.
Publicado: (2025)
SVFAP: Self-supervised Video Facial Affect Perceiver
por: Sun, Licai, et al.
Publicado: (2023)
por: Sun, Licai, et al.
Publicado: (2023)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
por: Foteinopoulou, Niki Maria, et al.
Publicado: (2023)
por: Foteinopoulou, Niki Maria, et al.
Publicado: (2023)
Can Large Language Models Capture Video Game Engagement?
por: Melhart, David, et al.
Publicado: (2025)
por: Melhart, David, et al.
Publicado: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
por: Verma, Arnav, et al.
Publicado: (2025)
por: Verma, Arnav, et al.
Publicado: (2025)
VideoMix: Aggregating How-To Videos for Task-Oriented Learning
por: Yang, Saelyne, et al.
Publicado: (2025)
por: Yang, Saelyne, et al.
Publicado: (2025)
A Review on Large Language Models for Visual Analytics
por: Agarwal, Navya Sonal, et al.
Publicado: (2025)
por: Agarwal, Navya Sonal, et al.
Publicado: (2025)
VideoA11y: Method and Dataset for Accessible Video Description
por: Li, Chaoyu, et al.
Publicado: (2025)
por: Li, Chaoyu, et al.
Publicado: (2025)
Vision Language Models as Values Detectors
por: Abbo, Giulio Antonio, et al.
Publicado: (2025)
por: Abbo, Giulio Antonio, et al.
Publicado: (2025)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
por: Huang, Yifei, et al.
Publicado: (2025)
por: Huang, Yifei, et al.
Publicado: (2025)
Do Vision Language Models Understand Human Engagement in Games?
por: Wang, Ziyi, et al.
Publicado: (2026)
por: Wang, Ziyi, et al.
Publicado: (2026)
ADAS-TO: A Large-Scale Multimodal Naturalistic Dataset and Empirical Characterization of Human Takeovers during ADAS Engagement
por: Wang, Yuhang, et al.
Publicado: (2026)
por: Wang, Yuhang, et al.
Publicado: (2026)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
por: Zhao, Yiming, et al.
Publicado: (2026)
por: Zhao, Yiming, et al.
Publicado: (2026)
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
por: Pan, Bo, et al.
Publicado: (2025)
por: Pan, Bo, et al.
Publicado: (2025)
Panda or not Panda? Understanding Adversarial Attacks with Interactive Visualization
por: You, Yuzhe, et al.
Publicado: (2023)
por: You, Yuzhe, et al.
Publicado: (2023)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
por: Yang, Jiaxi, et al.
Publicado: (2025)
por: Yang, Jiaxi, et al.
Publicado: (2025)
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
por: Ramachandran, Akhil, et al.
Publicado: (2026)
por: Ramachandran, Akhil, et al.
Publicado: (2026)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
por: Guo, Hao, et al.
Publicado: (2025)
por: Guo, Hao, et al.
Publicado: (2025)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
por: Liu, Can, et al.
Publicado: (2025)
por: Liu, Can, et al.
Publicado: (2025)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
por: Nowicki, Filip, et al.
Publicado: (2026)
por: Nowicki, Filip, et al.
Publicado: (2026)
Assessing Medical Training Skills via Eye and Head Movements
por: Latifzadeh, Kayhan, et al.
Publicado: (2025)
por: Latifzadeh, Kayhan, et al.
Publicado: (2025)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
por: Chen, Tianrun, et al.
Publicado: (2024)
por: Chen, Tianrun, et al.
Publicado: (2024)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
por: Xenos, Alexandros, et al.
Publicado: (2024)
por: Xenos, Alexandros, et al.
Publicado: (2024)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
por: Li, Chentao, et al.
Publicado: (2026)
por: Li, Chentao, et al.
Publicado: (2026)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
por: Nadeem, Asmar, et al.
Publicado: (2024)
por: Nadeem, Asmar, et al.
Publicado: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
por: Huh, Mina, et al.
Publicado: (2025)
por: Huh, Mina, et al.
Publicado: (2025)
Analyzing Swimming Performance Using Drone Captured Aerial Videos
por: Tran, Thu, et al.
Publicado: (2025)
por: Tran, Thu, et al.
Publicado: (2025)
Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
por: Eing, Lennart, et al.
Publicado: (2026)
por: Eing, Lennart, et al.
Publicado: (2026)
Ejemplares similares
-
MindCross: Fast New Subject Adaptation with Limited Data for Cross-subject Video Reconstruction from Brain Signals
por: Liu, Xuan-Hao, et al.
Publicado: (2025) -
Deep Neural Encoder-Decoder Model to Relate fMRI Brain Activity with Naturalistic Stimuli
por: David, Florian, et al.
Publicado: (2025) -
Adaptive Modality Balanced Online Knowledge Distillation for Brain-Eye-Computer based Dim Object Detection
por: Li, Zixing, et al.
Publicado: (2024) -
ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding
por: Liu, Minxu, et al.
Publicado: (2025) -
Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
por: Zhang, Guowei, et al.
Publicado: (2025)