Visual moral inference and communication
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Warren, Ramezani, Aida, Xu, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
por: Jing, Liqi, et al.
Publicado: (2026)
por: Jing, Liqi, et al.
Publicado: (2026)
Unsupervised Audio-Visual Segmentation with Modality Alignment
por: Bhosale, Swapnil, et al.
Publicado: (2024)
por: Bhosale, Swapnil, et al.
Publicado: (2024)
Beyond Frequency: Seeing Subtle Cues Through the Lens of Spatial Decomposition for Fine-Grained Visual Classification
por: Xu, Qin, et al.
Publicado: (2025)
por: Xu, Qin, et al.
Publicado: (2025)
MITracker: Multi-View Integration for Visual Object Tracking
por: Xu, Mengjie, et al.
Publicado: (2025)
por: Xu, Mengjie, et al.
Publicado: (2025)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
por: Dai, Ming, et al.
Publicado: (2024)
por: Dai, Ming, et al.
Publicado: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
por: Liu, Xiaolin, et al.
Publicado: (2026)
por: Liu, Xiaolin, et al.
Publicado: (2026)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
por: Chen, Pingyi, et al.
Publicado: (2024)
por: Chen, Pingyi, et al.
Publicado: (2024)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
por: Xu, Beining, et al.
Publicado: (2025)
por: Xu, Beining, et al.
Publicado: (2025)
Compensating Visual Insufficiency with Stratified Language Guidance for Long-Tail Class Incremental Learning
por: Wang, Xi, et al.
Publicado: (2026)
por: Wang, Xi, et al.
Publicado: (2026)
Learning Physical Dynamics for Object-centric Visual Prediction
por: Xu, Huilin, et al.
Publicado: (2024)
por: Xu, Huilin, et al.
Publicado: (2024)
Bridge then Begin Anew: Generating Target-relevant Intermediate Model for Source-free Visual Emotion Adaptation
por: Zhu, Jiankun, et al.
Publicado: (2024)
por: Zhu, Jiankun, et al.
Publicado: (2024)
DTL: Disentangled Transfer Learning for Visual Recognition
por: Fu, Minghao, et al.
Publicado: (2023)
por: Fu, Minghao, et al.
Publicado: (2023)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
por: Zhan, Yufei, et al.
Publicado: (2025)
por: Zhan, Yufei, et al.
Publicado: (2025)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
por: Zhu, Xuanyu, et al.
Publicado: (2026)
por: Zhu, Xuanyu, et al.
Publicado: (2026)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
por: Gao, Mingjian, et al.
Publicado: (2026)
por: Gao, Mingjian, et al.
Publicado: (2026)
DOGR: Towards Versatile Visual Document Grounding and Referring
por: Zhou, Yinan, et al.
Publicado: (2024)
por: Zhou, Yinan, et al.
Publicado: (2024)
Contextual inference from single objects in Vision-Language models
por: Vilas, Martina G., et al.
Publicado: (2026)
por: Vilas, Martina G., et al.
Publicado: (2026)
Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and Benchmarking
por: Tang, Zhangyong, et al.
Publicado: (2025)
por: Tang, Zhangyong, et al.
Publicado: (2025)
Exploring Task-Level Optimal Prompts for Visual In-Context Learning
por: Zhu, Yan, et al.
Publicado: (2025)
por: Zhu, Yan, et al.
Publicado: (2025)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
por: Yu, Yinfeng, et al.
Publicado: (2025)
por: Yu, Yinfeng, et al.
Publicado: (2025)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
por: Xu, Haoran, et al.
Publicado: (2026)
por: Xu, Haoran, et al.
Publicado: (2026)
Visual-ERM: Reward Modeling for Visual Equivalence
por: Liu, Ziyu, et al.
Publicado: (2026)
por: Liu, Ziyu, et al.
Publicado: (2026)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
por: Ding, Zijun, et al.
Publicado: (2025)
por: Ding, Zijun, et al.
Publicado: (2025)
Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D Environments
por: Yang, Xiao, et al.
Publicado: (2025)
por: Yang, Xiao, et al.
Publicado: (2025)
Mutual Information guided Visual Contrastive Learning
por: Chen, Hanyang, et al.
Publicado: (2025)
por: Chen, Hanyang, et al.
Publicado: (2025)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
por: Suo, Yucheng, et al.
Publicado: (2025)
por: Suo, Yucheng, et al.
Publicado: (2025)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
por: Zhan, Yufei, et al.
Publicado: (2024)
por: Zhan, Yufei, et al.
Publicado: (2024)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
por: Zheng, Shurong, et al.
Publicado: (2026)
por: Zheng, Shurong, et al.
Publicado: (2026)
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes
por: Chen, Jingru, et al.
Publicado: (2026)
por: Chen, Jingru, et al.
Publicado: (2026)
Visual Space Optimization for Zero-shot Learning
por: Wang, Xinsheng, et al.
Publicado: (2019)
por: Wang, Xinsheng, et al.
Publicado: (2019)
Adversarial Error Correction for Visual Autoregressive Generation
por: Bi, Ligong, et al.
Publicado: (2026)
por: Bi, Ligong, et al.
Publicado: (2026)
Poivre: Self-Refining Visual Pointing with Reinforcement Learning
por: Yang, Wenjie, et al.
Publicado: (2025)
por: Yang, Wenjie, et al.
Publicado: (2025)
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
por: Lu, Wenting, et al.
Publicado: (2026)
por: Lu, Wenting, et al.
Publicado: (2026)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
por: Yang, Ziteng, et al.
Publicado: (2025)
por: Yang, Ziteng, et al.
Publicado: (2025)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
por: Xi, Suyang, et al.
Publicado: (2026)
por: Xi, Suyang, et al.
Publicado: (2026)
Breaking the accuracy-resource dilemma: a lightweight adaptive video inference enhancement
por: Ma, Wei, et al.
Publicado: (2026)
por: Ma, Wei, et al.
Publicado: (2026)
MemoNav: Working Memory Model for Visual Navigation
por: Li, Hongxin, et al.
Publicado: (2024)
por: Li, Hongxin, et al.
Publicado: (2024)
Dual Latent Memory for Visual Multi-agent System
por: Yu, Xinlei, et al.
Publicado: (2026)
por: Yu, Xinlei, et al.
Publicado: (2026)
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
por: Chen, Weiming, et al.
Publicado: (2025)
por: Chen, Weiming, et al.
Publicado: (2025)
An adversarial feature learning based semantic communication method for Human 3D Reconstruction
por: Liu, Shaojiang, et al.
Publicado: (2024)
por: Liu, Shaojiang, et al.
Publicado: (2024)
Ejemplares similares
-
StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
por: Jing, Liqi, et al.
Publicado: (2026) -
Unsupervised Audio-Visual Segmentation with Modality Alignment
por: Bhosale, Swapnil, et al.
Publicado: (2024) -
Beyond Frequency: Seeing Subtle Cues Through the Lens of Spatial Decomposition for Fine-Grained Visual Classification
por: Xu, Qin, et al.
Publicado: (2025) -
MITracker: Multi-View Integration for Visual Object Tracking
por: Xu, Mengjie, et al.
Publicado: (2025) -
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
por: Dai, Ming, et al.
Publicado: (2024)