WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
Fuente:
arXiv
Saved in:
| Main Authors: | Cherian, Anoop, Doyle, River, Ben-Dov, Eyal, Lohit, Suhas, Peng, Kuan-Chuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
by: Xiang, Xinhao, et al.
Published: (2025)
by: Xiang, Xinhao, et al.
Published: (2025)
Multimodal 3D Object Detection on Unseen Domains
by: Hegde, Deepti, et al.
Published: (2024)
by: Hegde, Deepti, et al.
Published: (2024)
Auto-Vocabulary 3D Object Detection
by: Zhang, Haomeng, et al.
Published: (2025)
by: Zhang, Haomeng, et al.
Published: (2025)
Equivariant Spatio-Temporal Self-Supervision for LiDAR Object Detection
by: Hegde, Deepti, et al.
Published: (2024)
by: Hegde, Deepti, et al.
Published: (2024)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
by: Li, Danrui, et al.
Published: (2026)
by: Li, Danrui, et al.
Published: (2026)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
by: Ni, Haomiao, et al.
Published: (2024)
by: Ni, Haomiao, et al.
Published: (2024)
Multimodal Diffusion Bridge with Attention-Based SAR Fusion for Satellite Image Cloud Removal
by: Hu, Yuyang, et al.
Published: (2025)
by: Hu, Yuyang, et al.
Published: (2025)
Improving Open-World Object Localization by Discovering Background
by: Singh, Ashish, et al.
Published: (2025)
by: Singh, Ashish, et al.
Published: (2025)
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
by: Kogashi, Kaen, et al.
Published: (2025)
by: Kogashi, Kaen, et al.
Published: (2025)
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models
by: Mumcu, Furkan, et al.
Published: (2026)
by: Mumcu, Furkan, et al.
Published: (2026)
Time-Series U-Net with Recurrence for Noise-Robust Imaging Photoplethysmography
by: Shenoy, Vineet R., et al.
Published: (2025)
by: Shenoy, Vineet R., et al.
Published: (2025)
Joint Training of Image Generator and Detector for Road Defect Detection
by: Peng, Kuan-Chuan
Published: (2025)
by: Peng, Kuan-Chuan
Published: (2025)
FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
by: Sawada, Naoko, et al.
Published: (2025)
by: Sawada, Naoko, et al.
Published: (2025)
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
by: Zhang, Jiahao, et al.
Published: (2023)
by: Zhang, Jiahao, et al.
Published: (2023)
LLM-Guided Agentic Object Detection for Open-World Understanding
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
ComplexVAD: Detecting Interaction Anomalies in Video
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
Programmatic Video Prediction Using Large Language Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Recovering Pulse Waves from Video Using Deep Unrolling and Deep Equilibrium Models
by: Shenoy, Vineet R, et al.
Published: (2025)
by: Shenoy, Vineet R, et al.
Published: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
by: Zeng, Ziyun, et al.
Published: (2025)
by: Zeng, Ziyun, et al.
Published: (2025)
Towards Zero-shot 3D Anomaly Localization
by: Wang, Yizhou, et al.
Published: (2024)
by: Wang, Yizhou, et al.
Published: (2024)
PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
by: Hou, Xiaolu, et al.
Published: (2025)
by: Hou, Xiaolu, et al.
Published: (2025)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Long-Tailed Anomaly Detection with Learnable Class Names
by: Ho, Chih-Hui, et al.
Published: (2024)
by: Ho, Chih-Hui, et al.
Published: (2024)
Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts
by: Yang, Chiao-An, et al.
Published: (2025)
by: Yang, Chiao-An, et al.
Published: (2025)
DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection
by: Qi, Hongyuan, et al.
Published: (2026)
by: Qi, Hongyuan, et al.
Published: (2026)
Memory-Distilled Selection for Noise-Robust Anomaly Detection
by: Safarov, Sirojbek, et al.
Published: (2026)
by: Safarov, Sirojbek, et al.
Published: (2026)
Enhancement-Driven Pretraining for Robust Fingerprint Representation Learning
by: Gavas, Ekta, et al.
Published: (2024)
by: Gavas, Ekta, et al.
Published: (2024)
A Robust Image Forensic Framework Utilizing Multi-Colorspace Enriched Vision Transformer for Distinguishing Natural and Computer-Generated Images
by: Gangan, Manjary P., et al.
Published: (2023)
by: Gangan, Manjary P., et al.
Published: (2023)
PF3Det: A Prompted Foundation Feature Assisted Visual LiDAR 3D Detector
by: Li, Kaidong, et al.
Published: (2025)
by: Li, Kaidong, et al.
Published: (2025)
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
by: Jiang, Yiwen, et al.
Published: (2025)
by: Jiang, Yiwen, et al.
Published: (2025)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
by: Ni, Yao, et al.
Published: (2025)
by: Ni, Yao, et al.
Published: (2025)
SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction
by: Fu, Haoxiang, et al.
Published: (2026)
by: Fu, Haoxiang, et al.
Published: (2026)
Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
Similar Items
-
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
by: Cherian, Anoop, et al.
Published: (2024) -
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
by: Xiang, Xinhao, et al.
Published: (2025) -
Multimodal 3D Object Detection on Unseen Domains
by: Hegde, Deepti, et al.
Published: (2024) -
Auto-Vocabulary 3D Object Detection
by: Zhang, Haomeng, et al.
Published: (2025) -
Equivariant Spatio-Temporal Self-Supervision for LiDAR Object Detection
by: Hegde, Deepti, et al.
Published: (2024)