Guardado en:
| Autores principales: | Su, Yiyang, Liu, Xiaoming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.05708 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
por: Zhu, Jie, et al.
Publicado: (2026)
por: Zhu, Jie, et al.
Publicado: (2026)
HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
por: Su, Yiyang, et al.
Publicado: (2025)
por: Su, Yiyang, et al.
Publicado: (2025)
KeyPoint Relative Position Encoding for Face Recognition
por: Kim, Minchul, et al.
Publicado: (2024)
por: Kim, Minchul, et al.
Publicado: (2024)
SapiensID: Foundation for Human Recognition
por: Kim, Minchul, et al.
Publicado: (2025)
por: Kim, Minchul, et al.
Publicado: (2025)
Open-Set Biometrics: Beyond Good Closed-Set Models
por: Su, Yiyang, et al.
Publicado: (2024)
por: Su, Yiyang, et al.
Publicado: (2024)
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
por: Wang, Yikun, et al.
Publicado: (2025)
por: Wang, Yikun, et al.
Publicado: (2025)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
por: Deng, Xiang, et al.
Publicado: (2026)
por: Deng, Xiang, et al.
Publicado: (2026)
FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
por: Zhu, Jie, et al.
Publicado: (2026)
por: Zhu, Jie, et al.
Publicado: (2026)
A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
por: Zhu, Jie, et al.
Publicado: (2025)
por: Zhu, Jie, et al.
Publicado: (2025)
Learning to Wander: Improving the Global Image Geolocation Ability of LMMs via Actionable Reasoning
por: Zheng, Yushuo, et al.
Publicado: (2026)
por: Zheng, Yushuo, et al.
Publicado: (2026)
LocalScore: Local Density-Aware Similarity Scoring for Biometrics
por: Su, Yiyang, et al.
Publicado: (2026)
por: Su, Yiyang, et al.
Publicado: (2026)
Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning
por: Qi, Mengshi, et al.
Publicado: (2025)
por: Qi, Mengshi, et al.
Publicado: (2025)
Statewide Visual Geolocalization in the Wild
por: Fervers, Florian, et al.
Publicado: (2024)
por: Fervers, Florian, et al.
Publicado: (2024)
Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework
por: Song, Zirui, et al.
Publicado: (2025)
por: Song, Zirui, et al.
Publicado: (2025)
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
por: Vyas, Apoorv, et al.
Publicado: (2025)
por: Vyas, Apoorv, et al.
Publicado: (2025)
Audiovisual Masked Autoencoders
por: Georgescu, Mariana-Iuliana, et al.
Publicado: (2022)
por: Georgescu, Mariana-Iuliana, et al.
Publicado: (2022)
GeoRC: A Benchmark for Geolocation Reasoning Chains
por: Talreja, Mohit, et al.
Publicado: (2026)
por: Talreja, Mohit, et al.
Publicado: (2026)
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
por: Chaubey, Ashutosh, et al.
Publicado: (2026)
por: Chaubey, Ashutosh, et al.
Publicado: (2026)
Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
por: Hashmi, Ammarah, et al.
Publicado: (2024)
por: Hashmi, Ammarah, et al.
Publicado: (2024)
Zwitscherkasten -- DIY Audiovisual bird monitoring
por: Blum, Dominik, et al.
Publicado: (2026)
por: Blum, Dominik, et al.
Publicado: (2026)
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector
por: Guo, Xiao, et al.
Publicado: (2025)
por: Guo, Xiao, et al.
Publicado: (2025)
PIGEON: Predicting Image Geolocations
por: Haas, Lukas, et al.
Publicado: (2023)
por: Haas, Lukas, et al.
Publicado: (2023)
GaGA: Towards Interactive Global Geolocation Assistant
por: Dou, Zhiyang, et al.
Publicado: (2024)
por: Dou, Zhiyang, et al.
Publicado: (2024)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
por: Gou, Yunhao, et al.
Publicado: (2025)
por: Gou, Yunhao, et al.
Publicado: (2025)
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
por: Liu, Yuqi, et al.
Publicado: (2025)
por: Liu, Yuqi, et al.
Publicado: (2025)
GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization
por: Jia, Pengyue, et al.
Publicado: (2026)
por: Jia, Pengyue, et al.
Publicado: (2026)
LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild
por: Wang, Zhiqiang, et al.
Publicado: (2024)
por: Wang, Zhiqiang, et al.
Publicado: (2024)
GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
por: Jia, Pengyue, et al.
Publicado: (2025)
por: Jia, Pengyue, et al.
Publicado: (2025)
Granular Privacy Control for Geolocation with Vision Language Models
por: Mendes, Ethan, et al.
Publicado: (2024)
por: Mendes, Ethan, et al.
Publicado: (2024)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
por: Chen, Xinlong, et al.
Publicado: (2025)
por: Chen, Xinlong, et al.
Publicado: (2025)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
por: Pillai, Manu S, et al.
Publicado: (2024)
por: Pillai, Manu S, et al.
Publicado: (2024)
Frequency-guided Multi-level Reasoning for Scene Graph Generation in Video
por: Li, Chenxing, et al.
Publicado: (2026)
por: Li, Chenxing, et al.
Publicado: (2026)
Referee: Reference-aware Audiovisual Deepfake Detection
por: Boo, Hyemin, et al.
Publicado: (2025)
por: Boo, Hyemin, et al.
Publicado: (2025)
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
por: Gadi, Hari Krishna, et al.
Publicado: (2026)
por: Gadi, Hari Krishna, et al.
Publicado: (2026)
Self-supervised Audiovisual Representation Learning for Remote Sensing Data
por: Heidler, Konrad, et al.
Publicado: (2021)
por: Heidler, Konrad, et al.
Publicado: (2021)
X-Streamer: Unified Human World Modeling with Audiovisual Interaction
por: Xie, You, et al.
Publicado: (2025)
por: Xie, You, et al.
Publicado: (2025)
Image-Based Geolocation Using Large Vision-Language Models
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
por: Qian, Zhaofang, et al.
Publicado: (2025)
por: Qian, Zhaofang, et al.
Publicado: (2025)
Combi-CAM: A Novel Multi-Layer Approach for Explainable Image Geolocalization
por: Faget, David, et al.
Publicado: (2026)
por: Faget, David, et al.
Publicado: (2026)
Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models
por: Grainge, Oliver, et al.
Publicado: (2025)
por: Grainge, Oliver, et al.
Publicado: (2025)
Ejemplares similares
-
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
por: Zhu, Jie, et al.
Publicado: (2026) -
HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
por: Su, Yiyang, et al.
Publicado: (2025) -
KeyPoint Relative Position Encoding for Face Recognition
por: Kim, Minchul, et al.
Publicado: (2024) -
SapiensID: Foundation for Human Recognition
por: Kim, Minchul, et al.
Publicado: (2025) -
Open-Set Biometrics: Beyond Good Closed-Set Models
por: Su, Yiyang, et al.
Publicado: (2024)