LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Zongyou, Qu, Qiang, Zhang, Qian, Zhang, Nan, Chen, Xiaoming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
von: Yu, Zongyou, et al.
Veröffentlicht: (2024)
von: Yu, Zongyou, et al.
Veröffentlicht: (2024)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes without References
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
Self-supervised Photographic Image Layout Representation Learning
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
FCBoost-Net: A Generative Network for Synthesizing Multiple Collocated Outfits via Fashion Compatibility Boosting
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Rethinking Video with a Universal Event-Based Representation
von: Freeman, Andrew
Veröffentlicht: (2024)
von: Freeman, Andrew
Veröffentlicht: (2024)
LFACon: Introducing Anglewise Attention to No-Reference Quality Assessment in Light Field Space
von: Qu, Qiang, et al.
Veröffentlicht: (2023)
von: Qu, Qiang, et al.
Veröffentlicht: (2023)
Enhancing Fake News Video Detection via LLM-Driven Creative Process Simulation
von: Bu, Yuyan, et al.
Veröffentlicht: (2025)
von: Bu, Yuyan, et al.
Veröffentlicht: (2025)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
MVBIND: Self-Supervised Music Recommendation For Videos Via Embedding Space Binding
von: Teng, Jiajie, et al.
Veröffentlicht: (2024)
von: Teng, Jiajie, et al.
Veröffentlicht: (2024)
Learning Brain Representation with Hierarchical Visual Embeddings
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
Efficient Self-Supervised Video Hashing with Selective State Spaces
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
von: Wang, Yihao, et al.
Veröffentlicht: (2024)
von: Wang, Yihao, et al.
Veröffentlicht: (2024)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
von: Lan, Xing, et al.
Veröffentlicht: (2024)
von: Lan, Xing, et al.
Veröffentlicht: (2024)
An Event-triggered System for Social Persuasion and Danger Alert in Elder Home Monitoring
von: Liu, Jun-Yi, et al.
Veröffentlicht: (2025)
von: Liu, Jun-Yi, et al.
Veröffentlicht: (2025)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
von: Chen, Nan, et al.
Veröffentlicht: (2024)
von: Chen, Nan, et al.
Veröffentlicht: (2024)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
von: Zhou, Shengli, et al.
Veröffentlicht: (2025)
von: Zhou, Shengli, et al.
Veröffentlicht: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Self-similarity Prior Distillation for Unsupervised Remote Physiological Measurement
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
A Simple Task-aware Contrastive Local Descriptor Selection Strategy for Few-shot Learning between inter class and intra class
von: Qiao, Qian, et al.
Veröffentlicht: (2024)
von: Qiao, Qian, et al.
Veröffentlicht: (2024)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
RETRO: REthinking Tactile Representation Learning with Material PriOrs
von: Xia, Weihao, et al.
Veröffentlicht: (2025)
von: Xia, Weihao, et al.
Veröffentlicht: (2025)
MSLIQA: Enhancing Learning Representations for Image Quality Assessment through Multi-Scale Learning
von: Avanaki, Nasim Jamshidi, et al.
Veröffentlicht: (2024)
von: Avanaki, Nasim Jamshidi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024) -
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
von: Yu, Zongyou, et al.
Veröffentlicht: (2024) -
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025) -
COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025) -
NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes without References
von: Qu, Qiang, et al.
Veröffentlicht: (2025)