Guardado en:
| Autores principales: | Ge, Junyao, Zhang, Xu, Zheng, Yang, Guo, Kaitai, Liang, Jimin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2408.14744 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
por: Lee, Kyuho, et al.
Publicado: (2025)
por: Lee, Kyuho, et al.
Publicado: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
por: Jian, Song, et al.
Publicado: (2025)
por: Jian, Song, et al.
Publicado: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
por: Su, Yuetong, et al.
Publicado: (2025)
por: Su, Yuetong, et al.
Publicado: (2025)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
por: Mehta, Vinit, et al.
Publicado: (2025)
por: Mehta, Vinit, et al.
Publicado: (2025)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
por: Wang, Chaoyi, et al.
Publicado: (2025)
por: Wang, Chaoyi, et al.
Publicado: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
por: Deng, Pei, et al.
Publicado: (2025)
por: Deng, Pei, et al.
Publicado: (2025)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
por: Wu, Jason, et al.
Publicado: (2026)
por: Wu, Jason, et al.
Publicado: (2026)
Embedding-Only Uplink for Onboard Retrieval Under Shift in Remote Sensing
por: Sim, Sangcheol
Publicado: (2026)
por: Sim, Sangcheol
Publicado: (2026)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
por: Chen, Jingkun, et al.
Publicado: (2025)
por: Chen, Jingkun, et al.
Publicado: (2025)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
por: Guo, Yijie, et al.
Publicado: (2025)
por: Guo, Yijie, et al.
Publicado: (2025)
From eye to AI: studying rodent social behavior in the era of machine Learning
por: Chindemi, Giuseppe, et al.
Publicado: (2025)
por: Chindemi, Giuseppe, et al.
Publicado: (2025)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
por: Hu, Yudong, et al.
Publicado: (2025)
por: Hu, Yudong, et al.
Publicado: (2025)
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
por: Liao, Guanghao, et al.
Publicado: (2026)
por: Liao, Guanghao, et al.
Publicado: (2026)
Semi supervised GAN for smart microscopy, fast and data efficient cell cycle classification
por: Manick, Rajeev, et al.
Publicado: (2026)
por: Manick, Rajeev, et al.
Publicado: (2026)
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
por: Syed, Shahram Najam, et al.
Publicado: (2025)
por: Syed, Shahram Najam, et al.
Publicado: (2025)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
por: Xiao, Jiasong, et al.
Publicado: (2026)
por: Xiao, Jiasong, et al.
Publicado: (2026)
From Dead Pixels to Editable Slides: Infographic Reconstruction into Native Google Slides via Vision-Language Region Understanding
por: Gonzalez, Leonardo
Publicado: (2026)
por: Gonzalez, Leonardo
Publicado: (2026)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
por: Chen, Yiteng, et al.
Publicado: (2025)
por: Chen, Yiteng, et al.
Publicado: (2025)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
por: Li, Xinqing, et al.
Publicado: (2025)
por: Li, Xinqing, et al.
Publicado: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
por: Pu, Qingwen, et al.
Publicado: (2026)
por: Pu, Qingwen, et al.
Publicado: (2026)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
por: Deka, Dawar Jyoti, et al.
Publicado: (2026)
por: Deka, Dawar Jyoti, et al.
Publicado: (2026)
Neighborhood Feature Pooling for Remote Sensing Image Classification
por: Nia, Fahimeh Orvati, et al.
Publicado: (2025)
por: Nia, Fahimeh Orvati, et al.
Publicado: (2025)
A Guide to Structureless Visual Localization
por: Panek, Vojtech, et al.
Publicado: (2025)
por: Panek, Vojtech, et al.
Publicado: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
por: Oliveira, Daniel, et al.
Publicado: (2026)
por: Oliveira, Daniel, et al.
Publicado: (2026)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
por: Wang, Yiming, et al.
Publicado: (2026)
por: Wang, Yiming, et al.
Publicado: (2026)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
por: Peng, Hongxing, et al.
Publicado: (2025)
por: Peng, Hongxing, et al.
Publicado: (2025)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
por: Panek, Vojtech, et al.
Publicado: (2024)
por: Panek, Vojtech, et al.
Publicado: (2024)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
por: Panek, Vojtech, et al.
Publicado: (2026)
por: Panek, Vojtech, et al.
Publicado: (2026)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
por: Bergkvist, Viktor, et al.
Publicado: (2026)
por: Bergkvist, Viktor, et al.
Publicado: (2026)
Exploring Surround-View Fisheye Camera 3D Object Detection
por: Li, Changcai, et al.
Publicado: (2025)
por: Li, Changcai, et al.
Publicado: (2025)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
por: Huang, Yian, et al.
Publicado: (2026)
por: Huang, Yian, et al.
Publicado: (2026)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
por: Karki, Siddhant, et al.
Publicado: (2025)
por: Karki, Siddhant, et al.
Publicado: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
por: Hou, Zhangcheng, et al.
Publicado: (2026)
por: Hou, Zhangcheng, et al.
Publicado: (2026)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
por: Hou, Zhiyi, et al.
Publicado: (2025)
por: Hou, Zhiyi, et al.
Publicado: (2025)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
por: Meng, Siyuan, et al.
Publicado: (2026)
por: Meng, Siyuan, et al.
Publicado: (2026)
Perceptual Flow Network for Visually Grounded Reasoning
por: Li, Yangfu, et al.
Publicado: (2026)
por: Li, Yangfu, et al.
Publicado: (2026)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
por: Chen, Yuangong, et al.
Publicado: (2026)
por: Chen, Yuangong, et al.
Publicado: (2026)
CoMatcher: Multi-View Collaborative Feature Matching
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
Ejemplares similares
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026) -
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
por: Lee, Kyuho, et al.
Publicado: (2025) -
A Vision-Language Model for Focal Liver Lesion Classification
por: Jian, Song, et al.
Publicado: (2025) -
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
por: Su, Yuetong, et al.
Publicado: (2025)