Enregistré dans:
| Auteurs principaux: | Ge, Junyao, Zhang, Xu, Zheng, Yang, Guo, Kaitai, Liang, Jimin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2408.14744 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
par: Raoufi, Behnam, et autres
Publié: (2025)
par: Raoufi, Behnam, et autres
Publié: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
par: Durrani, Hamza Ahmed, et autres
Publié: (2026)
par: Durrani, Hamza Ahmed, et autres
Publié: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
par: Lee, Kyuho, et autres
Publié: (2025)
par: Lee, Kyuho, et autres
Publié: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
par: Jian, Song, et autres
Publié: (2025)
par: Jian, Song, et autres
Publié: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
par: Su, Yuetong, et autres
Publié: (2025)
par: Su, Yuetong, et autres
Publié: (2025)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
par: Mehta, Vinit, et autres
Publié: (2025)
par: Mehta, Vinit, et autres
Publié: (2025)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
par: Wang, Chaoyi, et autres
Publié: (2025)
par: Wang, Chaoyi, et autres
Publié: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
par: Deng, Pei, et autres
Publié: (2025)
par: Deng, Pei, et autres
Publié: (2025)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
par: Wu, Jason, et autres
Publié: (2026)
par: Wu, Jason, et autres
Publié: (2026)
Embedding-Only Uplink for Onboard Retrieval Under Shift in Remote Sensing
par: Sim, Sangcheol
Publié: (2026)
par: Sim, Sangcheol
Publié: (2026)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
par: Chen, Jingkun, et autres
Publié: (2025)
par: Chen, Jingkun, et autres
Publié: (2025)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
par: Guo, Yijie, et autres
Publié: (2025)
par: Guo, Yijie, et autres
Publié: (2025)
From eye to AI: studying rodent social behavior in the era of machine Learning
par: Chindemi, Giuseppe, et autres
Publié: (2025)
par: Chindemi, Giuseppe, et autres
Publié: (2025)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
par: Hu, Yudong, et autres
Publié: (2025)
par: Hu, Yudong, et autres
Publié: (2025)
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
par: Liao, Guanghao, et autres
Publié: (2026)
par: Liao, Guanghao, et autres
Publié: (2026)
Semi supervised GAN for smart microscopy, fast and data efficient cell cycle classification
par: Manick, Rajeev, et autres
Publié: (2026)
par: Manick, Rajeev, et autres
Publié: (2026)
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
par: Syed, Shahram Najam, et autres
Publié: (2025)
par: Syed, Shahram Najam, et autres
Publié: (2025)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
par: Xiao, Jiasong, et autres
Publié: (2026)
par: Xiao, Jiasong, et autres
Publié: (2026)
From Dead Pixels to Editable Slides: Infographic Reconstruction into Native Google Slides via Vision-Language Region Understanding
par: Gonzalez, Leonardo
Publié: (2026)
par: Gonzalez, Leonardo
Publié: (2026)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
par: Chen, Yiteng, et autres
Publié: (2025)
par: Chen, Yiteng, et autres
Publié: (2025)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
par: Li, Xinqing, et autres
Publié: (2025)
par: Li, Xinqing, et autres
Publié: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
par: Pu, Qingwen, et autres
Publié: (2026)
par: Pu, Qingwen, et autres
Publié: (2026)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
par: Deka, Dawar Jyoti, et autres
Publié: (2026)
par: Deka, Dawar Jyoti, et autres
Publié: (2026)
Neighborhood Feature Pooling for Remote Sensing Image Classification
par: Nia, Fahimeh Orvati, et autres
Publié: (2025)
par: Nia, Fahimeh Orvati, et autres
Publié: (2025)
A Guide to Structureless Visual Localization
par: Panek, Vojtech, et autres
Publié: (2025)
par: Panek, Vojtech, et autres
Publié: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
par: Oliveira, Daniel, et autres
Publié: (2026)
par: Oliveira, Daniel, et autres
Publié: (2026)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
par: Wang, Yiming, et autres
Publié: (2026)
par: Wang, Yiming, et autres
Publié: (2026)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
par: Peng, Hongxing, et autres
Publié: (2025)
par: Peng, Hongxing, et autres
Publié: (2025)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
par: Panek, Vojtech, et autres
Publié: (2024)
par: Panek, Vojtech, et autres
Publié: (2024)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
par: Panek, Vojtech, et autres
Publié: (2026)
par: Panek, Vojtech, et autres
Publié: (2026)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
par: Bergkvist, Viktor, et autres
Publié: (2026)
par: Bergkvist, Viktor, et autres
Publié: (2026)
Exploring Surround-View Fisheye Camera 3D Object Detection
par: Li, Changcai, et autres
Publié: (2025)
par: Li, Changcai, et autres
Publié: (2025)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
par: Huang, Yian, et autres
Publié: (2026)
par: Huang, Yian, et autres
Publié: (2026)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
par: Karki, Siddhant, et autres
Publié: (2025)
par: Karki, Siddhant, et autres
Publié: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
par: Hou, Zhangcheng, et autres
Publié: (2026)
par: Hou, Zhangcheng, et autres
Publié: (2026)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
par: Hou, Zhiyi, et autres
Publié: (2025)
par: Hou, Zhiyi, et autres
Publié: (2025)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
par: Meng, Siyuan, et autres
Publié: (2026)
par: Meng, Siyuan, et autres
Publié: (2026)
Perceptual Flow Network for Visually Grounded Reasoning
par: Li, Yangfu, et autres
Publié: (2026)
par: Li, Yangfu, et autres
Publié: (2026)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
par: Chen, Yuangong, et autres
Publié: (2026)
par: Chen, Yuangong, et autres
Publié: (2026)
CoMatcher: Multi-View Collaborative Feature Matching
par: Zhang, Jintao, et autres
Publié: (2025)
par: Zhang, Jintao, et autres
Publié: (2025)
Documents similaires
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
par: Raoufi, Behnam, et autres
Publié: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
par: Durrani, Hamza Ahmed, et autres
Publié: (2026) -
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
par: Lee, Kyuho, et autres
Publié: (2025) -
A Vision-Language Model for Focal Liver Lesion Classification
par: Jian, Song, et autres
Publié: (2025) -
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
par: Su, Yuetong, et autres
Publié: (2025)