Guardado en:
| Autores principales: | Chen, Junwen, Wang, Yingcheng, Yanai, Keiji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2307.02291 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
por: Chen, Junwen, et al.
Publicado: (2025)
por: Chen, Junwen, et al.
Publicado: (2025)
PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
por: Xiong, Peilin, et al.
Publicado: (2025)
por: Xiong, Peilin, et al.
Publicado: (2025)
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
por: Xiong, Peilin, et al.
Publicado: (2026)
por: Xiong, Peilin, et al.
Publicado: (2026)
A DeNoising FPN With Transformer R-CNN for Tiny Object Detection
por: Liu, Hou-I, et al.
Publicado: (2024)
por: Liu, Hou-I, et al.
Publicado: (2024)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
por: Qiu, Yicheng, et al.
Publicado: (2026)
por: Qiu, Yicheng, et al.
Publicado: (2026)
Point-MF: One-step Point Cloud Generation from a Single Image via Mean Flows
por: Baba, Yuta, et al.
Publicado: (2026)
por: Baba, Yuta, et al.
Publicado: (2026)
U-DECN: End-to-End Underwater Object Detection ConvNet with Improved DeNoising Training
por: Liu, Zhuoyan, et al.
Publicado: (2024)
por: Liu, Zhuoyan, et al.
Publicado: (2024)
SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
por: Yuan, Honghui, et al.
Publicado: (2025)
por: Yuan, Honghui, et al.
Publicado: (2025)
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
por: Chen, Junwen, et al.
Publicado: (2025)
por: Chen, Junwen, et al.
Publicado: (2025)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
por: Vu, Kiet Dang, et al.
Publicado: (2026)
por: Vu, Kiet Dang, et al.
Publicado: (2026)
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training
por: Tao, Yiyi, et al.
Publicado: (2024)
por: Tao, Yiyi, et al.
Publicado: (2024)
Sample what you cant compress
por: Birodkar, Vighnesh, et al.
Publicado: (2024)
por: Birodkar, Vighnesh, et al.
Publicado: (2024)
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
por: Gomi, Keisuke, et al.
Publicado: (2026)
por: Gomi, Keisuke, et al.
Publicado: (2026)
From Image to Video, what do we need in multimodal LLMs?
por: Huang, Suyuan, et al.
Publicado: (2024)
por: Huang, Suyuan, et al.
Publicado: (2024)
Attend to what I say: Highlighting relevant content on slides
por: M, Megha Mariam K, et al.
Publicado: (2026)
por: M, Megha Mariam K, et al.
Publicado: (2026)
Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift
por: Chen, Lixian, et al.
Publicado: (2026)
por: Chen, Lixian, et al.
Publicado: (2026)
MultiAnimate: Pose-Guided Image Animation Made Extensible
por: Hu, Yingcheng, et al.
Publicado: (2026)
por: Hu, Yingcheng, et al.
Publicado: (2026)
3D Scene Graph Guided Vision-Language Pre-training
por: Liu, Hao, et al.
Publicado: (2024)
por: Liu, Hao, et al.
Publicado: (2024)
Anatomical Structure-Guided Medical Vision-Language Pre-training
por: Li, Qingqiu, et al.
Publicado: (2024)
por: Li, Qingqiu, et al.
Publicado: (2024)
CheXLearner: Text-Guided Fine-Grained Representation Learning for Progression Detection
por: Wang, Yuanzhuo, et al.
Publicado: (2025)
por: Wang, Yuanzhuo, et al.
Publicado: (2025)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
por: Tong, Yujia, et al.
Publicado: (2026)
por: Tong, Yujia, et al.
Publicado: (2026)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
por: Liu, Jihao, et al.
Publicado: (2024)
por: Liu, Jihao, et al.
Publicado: (2024)
Benford's law: what does it say on adversarial images?
por: Zago, João G., et al.
Publicado: (2021)
por: Zago, João G., et al.
Publicado: (2021)
Do not trust what you trust: Miscalibration in Semi-supervised Learning
por: Mishra, Shambhavi, et al.
Publicado: (2024)
por: Mishra, Shambhavi, et al.
Publicado: (2024)
SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models
por: Wang, Haobo, et al.
Publicado: (2026)
por: Wang, Haobo, et al.
Publicado: (2026)
Rethinking Cross-Dose PET Denoising: Mitigating Averaging Effects via Residual Noise Learning
por: Liu, Yichao, et al.
Publicado: (2026)
por: Liu, Yichao, et al.
Publicado: (2026)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
por: Fang, Yixiong, et al.
Publicado: (2024)
por: Fang, Yixiong, et al.
Publicado: (2024)
Giving each task what it needs -- leveraging structured sparsity for tailored multi-task learning
por: Upadhyay, Richa, et al.
Publicado: (2024)
por: Upadhyay, Richa, et al.
Publicado: (2024)
Attend what matters: Leveraging vision foundational models for breast cancer classification using mammograms
por: Sanghvi, Samyak, et al.
Publicado: (2026)
por: Sanghvi, Samyak, et al.
Publicado: (2026)
IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-training
por: Liu, Che, et al.
Publicado: (2023)
por: Liu, Che, et al.
Publicado: (2023)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
por: Xue, Wei, et al.
Publicado: (2026)
por: Xue, Wei, et al.
Publicado: (2026)
TeLL Me what you cant see
por: Cavasin, Saverio, et al.
Publicado: (2025)
por: Cavasin, Saverio, et al.
Publicado: (2025)
The Global-Local loop: what is missing in bridging the gap between geospatial data from numerous communities?
por: Mallet, Clément, et al.
Publicado: (2026)
por: Mallet, Clément, et al.
Publicado: (2026)
Flexible image analysis for law enforcement agencies with deep neural networks to determine: where, who and what
por: Bouma, Henri, et al.
Publicado: (2024)
por: Bouma, Henri, et al.
Publicado: (2024)
DM-FNet: Unified multimodal medical image fusion via diffusion process-trained encoder-decoder
por: He, Dan, et al.
Publicado: (2025)
por: He, Dan, et al.
Publicado: (2025)
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
por: Shen, Hui, et al.
Publicado: (2026)
por: Shen, Hui, et al.
Publicado: (2026)
Vision-language models for decoding provider attention during neonatal resuscitation
por: Parodi, Felipe, et al.
Publicado: (2024)
por: Parodi, Felipe, et al.
Publicado: (2024)
Understanding when spatial transformer networks do not support invariance, and what to do about it
por: Finnveden, Lukas, et al.
Publicado: (2020)
por: Finnveden, Lukas, et al.
Publicado: (2020)
Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
por: Che, Chengan, et al.
Publicado: (2026)
por: Che, Chengan, et al.
Publicado: (2026)
Vision Large Language Models Are Good Noise Handlers in Engagement Analysis
por: Vedernikov, Alexander, et al.
Publicado: (2025)
por: Vedernikov, Alexander, et al.
Publicado: (2025)
Ejemplares similares
-
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
por: Chen, Junwen, et al.
Publicado: (2025) -
PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
por: Xiong, Peilin, et al.
Publicado: (2025) -
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
por: Xiong, Peilin, et al.
Publicado: (2026) -
A DeNoising FPN With Transformer R-CNN for Tiny Object Detection
por: Liu, Hou-I, et al.
Publicado: (2024) -
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
por: Qiu, Yicheng, et al.
Publicado: (2026)