LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Zichao, Pang, Bowen, Huang, Xufeng, Ji, Hang, Zhan, Xin, Chen, Junbo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MV-DETR: Multi-modality indoor object detection by Multi-View DEtecton TRansformers
by: Dong, Zichao, et al.
Published: (2024)
by: Dong, Zichao, et al.
Published: (2024)
PeP: a Point enhanced Painting method for unified point cloud tasks
by: Dong, Zichao, et al.
Published: (2023)
by: Dong, Zichao, et al.
Published: (2023)
RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
by: Ji, Hang, et al.
Published: (2025)
by: Ji, Hang, et al.
Published: (2025)
Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
by: Wang, Haowei, et al.
Published: (2023)
by: Wang, Haowei, et al.
Published: (2023)
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
by: Ji, Jiayi, et al.
Published: (2023)
by: Ji, Jiayi, et al.
Published: (2023)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
by: Qi, Yukun, et al.
Published: (2026)
by: Qi, Yukun, et al.
Published: (2026)
Efficient Multi-modal Large Language Models via Visual Token Grouping
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
by: Wang, Jin, et al.
Published: (2025)
by: Wang, Jin, et al.
Published: (2025)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
by: Feng, X., et al.
Published: (2024)
by: Feng, X., et al.
Published: (2024)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
by: Wen, Jiahao, et al.
Published: (2025)
by: Wen, Jiahao, et al.
Published: (2025)
InfoMatch: Entropy Neural Estimation for Semi-Supervised Image Classification
by: Han, Qi, et al.
Published: (2024)
by: Han, Qi, et al.
Published: (2024)
Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models
by: Zhan, Yu-Wei, et al.
Published: (2023)
by: Zhan, Yu-Wei, et al.
Published: (2023)
InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
GeoMM: On Geodesic Perspective for Multi-modal Learning
by: Mei, Shibin, et al.
Published: (2025)
by: Mei, Shibin, et al.
Published: (2025)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
by: Liu, Weide, et al.
Published: (2024)
by: Liu, Weide, et al.
Published: (2024)
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
by: Luan, Bozhi, et al.
Published: (2025)
by: Luan, Bozhi, et al.
Published: (2025)
Fine-grained Context and Multi-modal Alignment for Freehand 3D Ultrasound Reconstruction
by: Yan, Zhongnuo, et al.
Published: (2024)
by: Yan, Zhongnuo, et al.
Published: (2024)
Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
FrontierNet: Learning Visual Cues to Explore
by: Sun, Boyang, et al.
Published: (2025)
by: Sun, Boyang, et al.
Published: (2025)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
InViC: Intent-aware Visual Cues for Medical Visual Question Answering
by: Wang, Zhisong, et al.
Published: (2026)
by: Wang, Zhisong, et al.
Published: (2026)
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID
by: Li, Jiaxuan, et al.
Published: (2026)
by: Li, Jiaxuan, et al.
Published: (2026)
Oracle Bone Inscriptions Multi-modal Dataset
by: Li, Bang, et al.
Published: (2024)
by: Li, Bang, et al.
Published: (2024)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
by: Dong, Sibo, et al.
Published: (2025)
by: Dong, Sibo, et al.
Published: (2025)
LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
by: Guo, Meng-Hao, et al.
Published: (2025)
by: Guo, Meng-Hao, et al.
Published: (2025)
Exploiting Polarized Material Cues for Robust Car Detection
by: Dong, Wen, et al.
Published: (2024)
by: Dong, Wen, et al.
Published: (2024)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
by: Li, Xudong, et al.
Published: (2025)
by: Li, Xudong, et al.
Published: (2025)
Pixel-Wise Contrastive Distillation
by: Huang, Junqiang, et al.
Published: (2022)
by: Huang, Junqiang, et al.
Published: (2022)
InfoAffect: Affective Annotations of Infographics in Information Spread
by: Fu, Zihang, et al.
Published: (2025)
by: Fu, Zihang, et al.
Published: (2025)
ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions
by: Chen, Dubing, et al.
Published: (2024)
by: Chen, Dubing, et al.
Published: (2024)
InfoNorm: Mutual Information Shaping of Normals for Sparse-View Reconstruction
by: Wang, Xulong, et al.
Published: (2024)
by: Wang, Xulong, et al.
Published: (2024)
InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Similar Items
-
MV-DETR: Multi-modality indoor object detection by Multi-View DEtecton TRansformers
by: Dong, Zichao, et al.
Published: (2024) -
PeP: a Point enhanced Painting method for unified point cloud tasks
by: Dong, Zichao, et al.
Published: (2023) -
RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
by: Ji, Hang, et al.
Published: (2025) -
Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
by: Wang, Haowei, et al.
Published: (2023) -
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
by: Ji, Jiayi, et al.
Published: (2023)