AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Bin, Wen, Shaoguo, Fan, Yang, Wu, Shunlong, Wang, Junjie, Li, Yulin, Zhao, Junzhi, Wang, Junle, Tian, Zhuotao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
by: Kang, Bin, et al.
Published: (2025)
by: Kang, Bin, et al.
Published: (2025)
DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
by: Li, Yulin, et al.
Published: (2025)
by: Li, Yulin, et al.
Published: (2025)
TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices
by: Wang, Junle, et al.
Published: (2026)
by: Wang, Junle, et al.
Published: (2026)
On Learning Closed-Loop Probabilistic Multi-Agent Simulator
by: Lu, Juanwu, et al.
Published: (2025)
by: Lu, Juanwu, et al.
Published: (2025)
Let the Agent Steer: Closed-Loop Ranking Optimization via Influence Exchange
by: Cheng, Yin, et al.
Published: (2026)
by: Cheng, Yin, et al.
Published: (2026)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
by: Xie, Tianxin, et al.
Published: (2025)
by: Xie, Tianxin, et al.
Published: (2025)
A Closed-Loop Multi-Agent Framework for Aerodynamics-Aware Automotive Styling Design
by: Jin, Xinyu, et al.
Published: (2025)
by: Jin, Xinyu, et al.
Published: (2025)
Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation
by: Liu, Yutong, et al.
Published: (2025)
by: Liu, Yutong, et al.
Published: (2025)
A Multi-Agent Human-LLM Collaborative Framework for Closed-Loop Scientific Literature Summarization
by: Jacobson, Maxwell J., et al.
Published: (2026)
by: Jacobson, Maxwell J., et al.
Published: (2026)
Smurfs: Multi-Agent System using Context-Efficient DFSDT for Tool Planning
by: Chen, Junzhi, et al.
Published: (2024)
by: Chen, Junzhi, et al.
Published: (2024)
Multi-Task Learning for Front-End Text Processing in TTS
by: Kang, Wonjune, et al.
Published: (2024)
by: Kang, Wonjune, et al.
Published: (2024)
Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval
by: Kang, Bin, et al.
Published: (2024)
by: Kang, Bin, et al.
Published: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
by: Ma, Ziyang, et al.
Published: (2023)
by: Ma, Ziyang, et al.
Published: (2023)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
by: Wu, Zhichao, et al.
Published: (2025)
by: Wu, Zhichao, et al.
Published: (2025)
Closed-Loop Vision-Language Planning for Multi-Agent Coordination
by: Li, Zhiyuan, et al.
Published: (2025)
by: Li, Zhiyuan, et al.
Published: (2025)
Governance-Aware Agent Telemetry for Closed-Loop Enforcement in Multi-Agent AI Systems
by: Pathak, Anshul, et al.
Published: (2026)
by: Pathak, Anshul, et al.
Published: (2026)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
by: Singh, Jaskaran, et al.
Published: (2025)
by: Singh, Jaskaran, et al.
Published: (2025)
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
by: Seki, Kentaro, et al.
Published: (2025)
by: Seki, Kentaro, et al.
Published: (2025)
STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices
by: Wang, Junle, et al.
Published: (2026)
by: Wang, Junle, et al.
Published: (2026)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
by: Liu, Yutong, et al.
Published: (2025)
by: Liu, Yutong, et al.
Published: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
by: Yin, Kang, et al.
Published: (2025)
by: Yin, Kang, et al.
Published: (2025)
EvoNash-MARL: A Closed-Loop Multi-Agent Reinforcement Learning Framework for Medium-Horizon Equity Allocation
by: Jia, Chongliu, et al.
Published: (2026)
by: Jia, Chongliu, et al.
Published: (2026)
Agentic Discovery: Closing the Loop with Cooperative Agents
by: Pauloski, J. Gregory, et al.
Published: (2025)
by: Pauloski, J. Gregory, et al.
Published: (2025)
PhysiAgent: An Embodied Agent Framework in Physical World
by: Wang, Zhihao, et al.
Published: (2025)
by: Wang, Zhihao, et al.
Published: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
by: Ren, Yong, et al.
Published: (2026)
by: Ren, Yong, et al.
Published: (2026)
Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation
by: Wang, Junjie, et al.
Published: (2026)
by: Wang, Junjie, et al.
Published: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
by: Han, Wooseok, et al.
Published: (2024)
by: Han, Wooseok, et al.
Published: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
by: Guan, Wenhao, et al.
Published: (2023)
by: Guan, Wenhao, et al.
Published: (2023)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
by: Guo, Hao-Han, et al.
Published: (2024)
by: Guo, Hao-Han, et al.
Published: (2024)
MunTTS: A Text-to-Speech System for Mundari
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
by: Qi, Xin, et al.
Published: (2024)
by: Qi, Xin, et al.
Published: (2024)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
by: Liang, Ziqi, et al.
Published: (2024)
by: Liang, Ziqi, et al.
Published: (2024)
Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
by: Zhang, Hongxiang, et al.
Published: (2026)
by: Zhang, Hongxiang, et al.
Published: (2026)
Certificate-Driven Closed-Loop Multi-Agent Path Finding with Inheritable Factorization
by: Li, Jiarui, et al.
Published: (2026)
by: Li, Jiarui, et al.
Published: (2026)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
by: Baik, Sangwon, et al.
Published: (2026)
by: Baik, Sangwon, et al.
Published: (2026)
Similar Items
-
CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
by: Kang, Bin, et al.
Published: (2025) -
DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
by: Wang, Junjie, et al.
Published: (2025) -
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
by: Li, Yulin, et al.
Published: (2025) -
TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices
by: Wang, Junle, et al.
Published: (2026) -
On Learning Closed-Loop Probabilistic Multi-Agent Simulator
by: Lu, Juanwu, et al.
Published: (2025)