Towards Flexible Visual Relationship Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Fangrui, Yang, Jianwei, Jiang, Huaizu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023)
by: Han, Zeyu, et al.
Published: (2023)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025)
by: Zhu, Fangrui, et al.
Published: (2025)
SNAP: Towards Segmenting Anything in Any Point Cloud
by: Gupta, Aniket, et al.
Published: (2025)
by: Gupta, Aniket, et al.
Published: (2025)
DCVNet: Dilated Cost Volume Networks for Fast Optical Flow
by: Jiang, Huaizu, et al.
Published: (2021)
by: Jiang, Huaizu, et al.
Published: (2021)
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
by: Ding, Tianye, et al.
Published: (2024)
by: Ding, Tianye, et al.
Published: (2024)
EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking
by: Zhu, Fangrui, et al.
Published: (2026)
by: Zhu, Fangrui, et al.
Published: (2026)
A Strong Baseline for Point Cloud Registration via Direct Superpoints Matching
by: Gupta, Aniket, et al.
Published: (2023)
by: Gupta, Aniket, et al.
Published: (2023)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026)
by: Goswami, Prajnan, et al.
Published: (2026)
NeuFlow: Real-time, High-accuracy Optical Flow Estimation on Robots Using Edge Devices
by: Zhang, Zhiyong, et al.
Published: (2024)
by: Zhang, Zhiyong, et al.
Published: (2024)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
by: Meng, Zichong, et al.
Published: (2024)
by: Meng, Zichong, et al.
Published: (2024)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
by: Nguyen, Hieu T., et al.
Published: (2024)
by: Nguyen, Hieu T., et al.
Published: (2024)
Absolute Coordinates Make Motion Generation Easy
by: Meng, Zichong, et al.
Published: (2025)
by: Meng, Zichong, et al.
Published: (2025)
LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
by: Liu, Chenying, et al.
Published: (2025)
by: Liu, Chenying, et al.
Published: (2025)
Towards Visual Query Segmentation in the Wild
by: Fan, Bing, et al.
Published: (2026)
by: Fan, Bing, et al.
Published: (2026)
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
by: Xie, Yiming, et al.
Published: (2024)
by: Xie, Yiming, et al.
Published: (2024)
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
by: Yao, Chun-Han, et al.
Published: (2025)
by: Yao, Chun-Han, et al.
Published: (2025)
SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
by: Tahboub, Hamza, et al.
Published: (2025)
by: Tahboub, Hamza, et al.
Published: (2025)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
by: Zhou, Jinxin, et al.
Published: (2025)
by: Zhou, Jinxin, et al.
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
SMooDi: Stylized Motion Diffusion Model
by: Zhong, Lei, et al.
Published: (2024)
by: Zhong, Lei, et al.
Published: (2024)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
by: Ding, Tianye, et al.
Published: (2025)
by: Ding, Tianye, et al.
Published: (2025)
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
NeuFlow v2: Push High-Efficiency Optical Flow To the Limit
by: Zhang, Zhiyong, et al.
Published: (2024)
by: Zhang, Zhiyong, et al.
Published: (2024)
Towards Flexible Evaluation for Generative Visual Question Answering
by: Ji, Huishan, et al.
Published: (2024)
by: Ji, Huishan, et al.
Published: (2024)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
by: Li, Yian, et al.
Published: (2026)
by: Li, Yian, et al.
Published: (2026)
FlexICL: A Flexible Visual In-context Learning Framework for Elbow and Wrist Ultrasound Segmentation
by: Zhou, Yuyue, et al.
Published: (2025)
by: Zhou, Yuyue, et al.
Published: (2025)
Ultrasound Nodule Segmentation Using Asymmetric Learning with Simple Clinical Annotation
by: Zhao, Xingyue, et al.
Published: (2024)
by: Zhao, Xingyue, et al.
Published: (2024)
Anatomically-guided masked autoencoder pre-training for aneurysm detection
by: Ceballos-Arroyo, Alberto Mario, et al.
Published: (2025)
by: Ceballos-Arroyo, Alberto Mario, et al.
Published: (2025)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
by: Meng, Lingchen, et al.
Published: (2024)
by: Meng, Lingchen, et al.
Published: (2024)
Towards Accurate Unified Anomaly Segmentation
by: Ma, Wenxin, et al.
Published: (2025)
by: Ma, Wenxin, et al.
Published: (2025)
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
Fast and Flexible Robustness Certificates for Semantic Segmentation
by: Massena, Thomas, et al.
Published: (2025)
by: Massena, Thomas, et al.
Published: (2025)
Implicit Counterfactual Learning for Audio-Visual Segmentation
by: Zha, Mingfeng, et al.
Published: (2025)
by: Zha, Mingfeng, et al.
Published: (2025)
Background Matters: A Cross-view Bidirectional Modeling Framework for Semi-supervised Medical Image Segmentation
by: Cao, Luyang, et al.
Published: (2025)
by: Cao, Luyang, et al.
Published: (2025)
Decoupled Competitive Framework for Semi-supervised Medical Image Segmentation
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
TokenSeg: Efficient 3D Medical Image Segmentation via Hierarchical Visual Token Compression
by: Zeng, Sen, et al.
Published: (2026)
by: Zeng, Sen, et al.
Published: (2026)
Similar Items
-
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023) -
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025) -
SNAP: Towards Segmenting Anything in Any Point Cloud
by: Gupta, Aniket, et al.
Published: (2025) -
DCVNet: Dilated Cost Volume Networks for Fast Optical Flow
by: Jiang, Huaizu, et al.
Published: (2021) -
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
by: Ding, Tianye, et al.
Published: (2024)