Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Zongyou, Qu, Qiang, Chen, Xiaoming, Wang, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
von: Sar, Ayan, et al.
Veröffentlicht: (2025)
von: Sar, Ayan, et al.
Veröffentlicht: (2025)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Audio-visual Event Localization on Portrait Mode Short Videos
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
Can We Edit Multimodal Large Language Models?
von: Cheng, Siyuan, et al.
Veröffentlicht: (2023)
von: Cheng, Siyuan, et al.
Veröffentlicht: (2023)
DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting
von: Lee, Seungjun, et al.
Veröffentlicht: (2025)
von: Lee, Seungjun, et al.
Veröffentlicht: (2025)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?
von: Li, Shuqing, et al.
Veröffentlicht: (2025)
von: Li, Shuqing, et al.
Veröffentlicht: (2025)
Efficient Low-Resolution Face Recognition via Bridge Distillation
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
von: Su, Qile, et al.
Veröffentlicht: (2025)
von: Su, Qile, et al.
Veröffentlicht: (2025)
NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes without References
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
von: Wang, Xiao, et al.
Veröffentlicht: (2023)
von: Wang, Xiao, et al.
Veröffentlicht: (2023)
Anti-Inpainting: A Proactive Defense Approach against Malicious Diffusion-based Inpainters under Unknown Conditions
von: Guo, Yimao, et al.
Veröffentlicht: (2025)
von: Guo, Yimao, et al.
Veröffentlicht: (2025)
Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models
von: Kim, Jihyun, et al.
Veröffentlicht: (2025)
von: Kim, Jihyun, et al.
Veröffentlicht: (2025)
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
Text-Only Data Synthesis for Vision Language Model Training
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
LMVD: A Large-Scale Multimodal Vlog Dataset for Depression Detection in the Wild
von: He, Lang, et al.
Veröffentlicht: (2024)
von: He, Lang, et al.
Veröffentlicht: (2024)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
von: Yu, Lijun, et al.
Veröffentlicht: (2023)
von: Yu, Lijun, et al.
Veröffentlicht: (2023)
PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
von: Yu, Zongyou, et al.
Veröffentlicht: (2025) -
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024) -
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025) -
E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
von: Qu, Qiang, et al.
Veröffentlicht: (2024) -
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2025)