RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xingqi, Zhang, Yujie, Gao, Shuyong, Li, Wenjie, Hong, Lingyi, Chen, Mingxi, Jiang, Kaixun, Fu, Jiyuan, Zhang, Wenqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
VideoPure: Diffusion-based Adversarial Purification for Video Recognition
by: Jiang, Kaixun, et al.
Published: (2025)
by: Jiang, Kaixun, et al.
Published: (2025)
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
by: Fu, Jiyuan, et al.
Published: (2026)
by: Fu, Jiyuan, et al.
Published: (2026)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
by: Fu, Jiyuan, et al.
Published: (2025)
by: Fu, Jiyuan, et al.
Published: (2025)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
ClickVOS: Click Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision
by: Hu, Junjie, et al.
Published: (2024)
by: Hu, Junjie, et al.
Published: (2024)
Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos
by: Feng, Yuang, et al.
Published: (2025)
by: Feng, Yuang, et al.
Published: (2025)
MSVCOD:A Large-Scale Multi-Scene Dataset for Video Camouflage Object Detection
by: Gao, Shuyong, et al.
Published: (2025)
by: Gao, Shuyong, et al.
Published: (2025)
General Compression Framework for Efficient Transformer Object Tracking
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection
by: Mok, Tsui Qin, et al.
Published: (2025)
by: Mok, Tsui Qin, et al.
Published: (2025)
Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization
by: Wang, Jiafeng, et al.
Published: (2022)
by: Wang, Jiafeng, et al.
Published: (2022)
Improving Adversarial Transferability with Neighbourhood Gradient Information
by: Guo, Haijing, et al.
Published: (2024)
by: Guo, Haijing, et al.
Published: (2024)
LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
by: Huang, Donghao, et al.
Published: (2026)
by: Huang, Donghao, et al.
Published: (2026)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
by: Huang, Xu, et al.
Published: (2025)
by: Huang, Xu, et al.
Published: (2025)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
by: Hong, Lingyi, et al.
Published: (2026)
by: Hong, Lingyi, et al.
Published: (2026)
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
by: Zhou, Yitong, et al.
Published: (2026)
by: Zhou, Yitong, et al.
Published: (2026)
Boosting Adversarial Transferability with Spatial Adversarial Alignment
by: Chen, Zhaoyu, et al.
Published: (2025)
by: Chen, Zhaoyu, et al.
Published: (2025)
Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
by: Jiang, Kaixun, et al.
Published: (2025)
by: Jiang, Kaixun, et al.
Published: (2025)
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
by: Jiang, Kaixun, et al.
Published: (2026)
by: Jiang, Kaixun, et al.
Published: (2026)
PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
by: Yan, Shilin, et al.
Published: (2023)
by: Yan, Shilin, et al.
Published: (2023)
OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
by: Li, Wanyun, et al.
Published: (2024)
by: Li, Wanyun, et al.
Published: (2024)
Delving into Decision-based Black-box Attacks on Semantic Segmentation
by: Chen, Zhaoyu, et al.
Published: (2024)
by: Chen, Zhaoyu, et al.
Published: (2024)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
by: Lu, Yifei, et al.
Published: (2025)
by: Lu, Yifei, et al.
Published: (2025)
OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
by: Liu, Yilong, et al.
Published: (2026)
by: Liu, Yilong, et al.
Published: (2026)
Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis
by: Song, Da, et al.
Published: (2026)
by: Song, Da, et al.
Published: (2026)
Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialog
by: Estienne, Lautaro, et al.
Published: (2025)
by: Estienne, Lautaro, et al.
Published: (2025)
Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval
by: Chen, Yanfei, et al.
Published: (2024)
by: Chen, Yanfei, et al.
Published: (2024)
Collaborative Reconstruction and Repair for Multi-class Industrial Anomaly Detection
by: Wang, Qishan, et al.
Published: (2025)
by: Wang, Qishan, et al.
Published: (2025)
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation
by: Zhu, Dongsheng, et al.
Published: (2025)
by: Zhu, Dongsheng, et al.
Published: (2025)
Similar Items
-
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
by: Fu, Jiyuan, et al.
Published: (2024) -
VideoPure: Diffusion-based Adversarial Purification for Video Recognition
by: Jiang, Kaixun, et al.
Published: (2025) -
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
by: Fu, Jiyuan, et al.
Published: (2026) -
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
by: Fu, Jiyuan, et al.
Published: (2024) -
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
by: Fu, Jiyuan, et al.
Published: (2025)