Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds
Fuente:
arXiv
Saved in:
| Main Authors: | Shaar, Eitan, Shaulov, Ariel, Chechik, Gal, Wolf, Lior |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compositional Video Generation via Inference-Time Guidance
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
by: Shaar, Eitan, et al.
Published: (2026)
by: Shaar, Eitan, et al.
Published: (2026)
Classifier-Guided Captioning Across Modalities
by: Shaulov, Ariel, et al.
Published: (2025)
by: Shaulov, Ariel, et al.
Published: (2025)
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
by: Shaulov, Ariel, et al.
Published: (2025)
by: Shaulov, Ariel, et al.
Published: (2025)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Training-Free Consistent Text-to-Image Generation
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Per-Query Visual Concept Learning
by: Malca, Ori, et al.
Published: (2025)
by: Malca, Ori, et al.
Published: (2025)
ConsiStyle: Style Diversity in Training-Free Consistent T2I Generation
by: Mazuz, Yohai, et al.
Published: (2025)
by: Mazuz, Yohai, et al.
Published: (2025)
SphereUFormer: A U-Shaped Transformer for Spherical 360 Perception
by: Benny, Yaniv, et al.
Published: (2024)
by: Benny, Yaniv, et al.
Published: (2024)
DiffUHaul: A Training-Free Method for Object Dragging in Images
by: Avrahami, Omri, et al.
Published: (2024)
by: Avrahami, Omri, et al.
Published: (2024)
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
by: Samuel, Dvir, et al.
Published: (2025)
by: Samuel, Dvir, et al.
Published: (2025)
Data-Driven Loss Functions for Inference-Time Optimization in Text-to-Image
by: Yiflach, Sapir Esther, et al.
Published: (2025)
by: Yiflach, Sapir Esther, et al.
Published: (2025)
Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation
by: Atzmon, Yuval, et al.
Published: (2024)
by: Atzmon, Yuval, et al.
Published: (2024)
IT$^3$: Idempotent Test-Time Training
by: Durasov, Nikita, et al.
Published: (2024)
by: Durasov, Nikita, et al.
Published: (2024)
Bringing Objects to Life: training-free 4D generation from 3D objects through view consistent noise
by: Rahamim, Ohad, et al.
Published: (2024)
by: Rahamim, Ohad, et al.
Published: (2024)
Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
by: Samuel, Dvir, et al.
Published: (2026)
by: Samuel, Dvir, et al.
Published: (2026)
Single Image Iterative Subject-driven Generation and Editing
by: Shpitzer, Yair, et al.
Published: (2025)
by: Shpitzer, Yair, et al.
Published: (2025)
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
by: Luo, Zhengyi, et al.
Published: (2025)
by: Luo, Zhengyi, et al.
Published: (2025)
Key-Locked Rank One Editing for Text-to-Image Personalization
by: Tewel, Yoad, et al.
Published: (2023)
by: Tewel, Yoad, et al.
Published: (2023)
Diffusion-Based Attention Warping for Consistent 3D Scene Editing
by: Gomel, Eyal, et al.
Published: (2024)
by: Gomel, Eyal, et al.
Published: (2024)
Policy Optimized Text-to-Image Pipeline Design
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
TriTex: Learning Texture from a Single Mesh via Triplane Semantic Features
by: Cohen-Bar, Dana, et al.
Published: (2025)
by: Cohen-Bar, Dana, et al.
Published: (2025)
LCM-Lookahead for Encoder-based Text-to-Image Personalization
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
by: Manor, Hila, et al.
Published: (2026)
by: Manor, Hila, et al.
Published: (2026)
Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models
by: Toker, Michael, et al.
Published: (2025)
by: Toker, Michael, et al.
Published: (2025)
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
by: Dahan, Aviad, et al.
Published: (2026)
by: Dahan, Aviad, et al.
Published: (2026)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
IlluSign: Illustrating Sign Language Videos by Leveraging the Attention Mechanism
by: Bruner, Janna, et al.
Published: (2025)
by: Bruner, Janna, et al.
Published: (2025)
A Meaningful Perturbation Metric for Evaluating Explainability Methods
by: Cohen, Danielle, et al.
Published: (2025)
by: Cohen, Danielle, et al.
Published: (2025)
Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors
by: Rahamim, Ohad, et al.
Published: (2024)
by: Rahamim, Ohad, et al.
Published: (2024)
Audio-Guided Visual Perception for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Towards Open-Vocabulary Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
EventVAD: Training-Free Event-Aware Video Anomaly Detection
by: Shao, Yihua, et al.
Published: (2025)
by: Shao, Yihua, et al.
Published: (2025)
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
by: Ahmadian, Mona, et al.
Published: (2025)
by: Ahmadian, Mona, et al.
Published: (2025)
LTX-2: Efficient Joint Audio-Visual Foundation Model
by: HaCohen, Yoav, et al.
Published: (2026)
by: HaCohen, Yoav, et al.
Published: (2026)
REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
by: Almog, Gal, et al.
Published: (2025)
by: Almog, Gal, et al.
Published: (2025)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
by: Rassin, Royi, et al.
Published: (2023)
by: Rassin, Royi, et al.
Published: (2023)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
Similar Items
-
Compositional Video Generation via Inference-Time Guidance
by: Shaulov, Ariel, et al.
Published: (2026) -
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
by: Shaulov, Ariel, et al.
Published: (2026) -
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
by: Shaar, Eitan, et al.
Published: (2026) -
Classifier-Guided Captioning Across Modalities
by: Shaulov, Ariel, et al.
Published: (2025) -
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
by: Shaulov, Ariel, et al.
Published: (2025)