FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xuehai, Zheng, Jian, Fang, Jacob Zhiyuan, Piramuthu, Robinson, Bansal, Mohit, Ordonez, Vicente, Sigurdsson, Gunnar A, Peng, Nanyun, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InSTA: Towards Internet-Scale Training For Agents
by: Trabucco, Brandon, et al.
Published: (2025)
by: Trabucco, Brandon, et al.
Published: (2025)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
by: Zhang, Junyi, et al.
Published: (2025)
by: Zhang, Junyi, et al.
Published: (2025)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025)
by: Wadhawan, Rohan, et al.
Published: (2025)
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025)
by: Zhang, Zhiyuan, et al.
Published: (2025)
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert Controllers
by: Li, Sijia, et al.
Published: (2023)
by: Li, Sijia, et al.
Published: (2023)
REAL Sampling: Boosting Factuality and Diversity of Open-Ended Generation via Asymptotic Entropy
by: Chang, Haw-Shiuan, et al.
Published: (2024)
by: Chang, Haw-Shiuan, et al.
Published: (2024)
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM
by: Chang, Haw-Shiuan, et al.
Published: (2024)
by: Chang, Haw-Shiuan, et al.
Published: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
by: Yu, Shoubin, et al.
Published: (2024)
by: Yu, Shoubin, et al.
Published: (2024)
FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
by: Yuan, Shengming, et al.
Published: (2025)
by: Yuan, Shengming, et al.
Published: (2025)
FlexControl: Computation-Aware ControlNet with Differentiable Router for Text-to-Image Generation
by: Fang, Zheng, et al.
Published: (2025)
by: Fang, Zheng, et al.
Published: (2025)
FetalFlex: Anatomy-Guided Diffusion Model for Flexible Control on Fetal Ultrasound Image Synthesis
by: Duan, Yaofei, et al.
Published: (2025)
by: Duan, Yaofei, et al.
Published: (2025)
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
by: Sheng, Mingzhi, et al.
Published: (2026)
by: Sheng, Mingzhi, et al.
Published: (2026)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents
by: Pham, Quang Hieu, et al.
Published: (2026)
by: Pham, Quang Hieu, et al.
Published: (2026)
FlexGen: Flexible Multi-View Generation from Text and Image Inputs
by: Xu, Xinli, et al.
Published: (2024)
by: Xu, Xinli, et al.
Published: (2024)
FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction
by: Xu, Muhao, et al.
Published: (2024)
by: Xu, Muhao, et al.
Published: (2024)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA
by: Yan, Qianqi, et al.
Published: (2024)
by: Yan, Qianqi, et al.
Published: (2024)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
FlexHB: a More Efficient and Flexible Framework for Hyperparameter Optimization
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
by: Chen, Jiahao, et al.
Published: (2025)
by: Chen, Jiahao, et al.
Published: (2025)
FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models
by: Pirchert, Annemette Brok, et al.
Published: (2026)
by: Pirchert, Annemette Brok, et al.
Published: (2026)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
by: Ma, Linhan, et al.
Published: (2025)
by: Ma, Linhan, et al.
Published: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
by: Yun, Sukwon, et al.
Published: (2024)
by: Yun, Sukwon, et al.
Published: (2024)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
FlexEdit: Flexible and Controllable Diffusion-based Object-centric Image Editing
by: Nguyen, Trong-Tung, et al.
Published: (2024)
by: Nguyen, Trong-Tung, et al.
Published: (2024)
Flex-GAD : Flexible Graph Anomaly Detection
by: Chakraborty, Apu, et al.
Published: (2025)
by: Chakraborty, Apu, et al.
Published: (2025)
FlexPara: Flexible Neural Surface Parameterization
by: Zhao, Yuming, et al.
Published: (2025)
by: Zhao, Yuming, et al.
Published: (2025)
FlexOlmo: Open Language Models for Flexible Data Use
by: Shi, Weijia, et al.
Published: (2025)
by: Shi, Weijia, et al.
Published: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
by: Liu, Yinhong, et al.
Published: (2025)
by: Liu, Yinhong, et al.
Published: (2025)
FlexICL: A Flexible Visual In-context Learning Framework for Elbow and Wrist Ultrasound Segmentation
by: Zhou, Yuyue, et al.
Published: (2025)
by: Zhou, Yuyue, et al.
Published: (2025)
Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
by: Maharana, Adyasha, et al.
Published: (2024)
by: Maharana, Adyasha, et al.
Published: (2024)
Similar Items
-
InSTA: Towards Internet-Scale Training For Agents
by: Trabucco, Brandon, et al.
Published: (2025) -
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
by: Zhang, Junyi, et al.
Published: (2025) -
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024) -
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025) -
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025)