FIRM: Flexible Interactive Reflection reMoval
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xiao, Jiang, Xudong, Tao, Yunkang, Lei, Zhen, Li, Qing, Lei, Chenyang, Zhang, Zhaoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality
by: Lei, Chenyang, et al.
Published: (2024)
by: Lei, Chenyang, et al.
Published: (2024)
SimMAT: Exploring Transferability from Vision Foundation Models to Any Image Modality
by: Lei, Chenyang, et al.
Published: (2024)
by: Lei, Chenyang, et al.
Published: (2024)
Generative Active Learning for Image Synthesis Personalization
by: Zhang, Xulu, et al.
Published: (2024)
by: Zhang, Xulu, et al.
Published: (2024)
Compositional Inversion for Stable Diffusion Models
by: Zhang, Xulu, et al.
Published: (2023)
by: Zhang, Xulu, et al.
Published: (2023)
General Geometry-aware Weakly Supervised 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2024)
by: Zhang, Guowen, et al.
Published: (2024)
Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
by: Wang, Pengfei, et al.
Published: (2024)
by: Wang, Pengfei, et al.
Published: (2024)
Top-Down Guidance for Learning Object-Centric Representations
by: Zou, Junhong, et al.
Published: (2024)
by: Zou, Junhong, et al.
Published: (2024)
Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
by: Chen, Zuyao, et al.
Published: (2023)
by: Chen, Zuyao, et al.
Published: (2023)
GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives
by: Chen, Zuyao, et al.
Published: (2023)
by: Chen, Zuyao, et al.
Published: (2023)
Generating on Generated: An Approach Towards Self-Evolving Diffusion Models
by: Zhang, Xulu, et al.
Published: (2025)
by: Zhang, Xulu, et al.
Published: (2025)
Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2024)
by: Zhang, Guowen, et al.
Published: (2024)
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
by: Li, Hongxin, et al.
Published: (2026)
by: Li, Hongxin, et al.
Published: (2026)
Robust Depth Enhancement via Polarization Prompt Fusion Tuning
by: Ikemura, Kei, et al.
Published: (2024)
by: Ikemura, Kei, et al.
Published: (2024)
A Survey on Personalized Content Synthesis with Diffusion Models
by: Zhang, Xulu, et al.
Published: (2024)
by: Zhang, Xulu, et al.
Published: (2024)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
Revisiting Marr in Face: The Building of 2D--2.5D--3D Representations in Deep Neural Networks
by: Zhu, Xiangyu, et al.
Published: (2024)
by: Zhu, Xiangyu, et al.
Published: (2024)
A Masked Reverse Knowledge Distillation Method Incorporating Global and Local Information for Image Anomaly Detection
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Prototypical Learning Guided Context-Aware Segmentation Network for Few-Shot Anomaly Detection
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
by: Yu, Chang, et al.
Published: (2024)
by: Yu, Chang, et al.
Published: (2024)
SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decomposition
by: Hu, Xu, et al.
Published: (2024)
by: Hu, Xu, et al.
Published: (2024)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
DC-Gaussian: Improving 3D Gaussian Splatting for Reflective Dash Cam Videos
by: Wang, Linhan, et al.
Published: (2024)
by: Wang, Linhan, et al.
Published: (2024)
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
by: Zhou, Jiazhou, et al.
Published: (2026)
by: Zhou, Jiazhou, et al.
Published: (2026)
Improving Large Vision-Language Models' Understanding for Flow Field Data
by: Zhang, Xiaomei, et al.
Published: (2025)
by: Zhang, Xiaomei, et al.
Published: (2025)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly Detection
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Attention Fusion Reverse Distillation for Multi-Lighting Image Anomaly Detection
by: Zhang, Yiheng, et al.
Published: (2024)
by: Zhang, Yiheng, et al.
Published: (2024)
FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering
by: Zhou, Jingqiu, et al.
Published: (2025)
by: Zhou, Jingqiu, et al.
Published: (2025)
Automatic Controllable Colorization via Imagination
by: Cong, Xiaoyan, et al.
Published: (2024)
by: Cong, Xiaoyan, et al.
Published: (2024)
ADNet: A Large-Scale and Extensible Multi-Domain Benchmark for Anomaly Detection Across 380 Real-World Categories
by: Ling, Hai, et al.
Published: (2025)
by: Ling, Hai, et al.
Published: (2025)
RAD: A Comprehensive Dataset for Benchmarking the Robustness of Image Anomaly Detection
by: Cheng, Yuqi, et al.
Published: (2024)
by: Cheng, Yuqi, et al.
Published: (2024)
Leveraging Local Patch Alignment to Seam-cutting for Large Parallax Image Stitching
by: Liao, Tianli, et al.
Published: (2023)
by: Liao, Tianli, et al.
Published: (2023)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
Towards Active Real-to-Twin Inspection: A New Paradigm for Zero-Shot Anomaly Detection
by: Liu, Jiaxuan, et al.
Published: (2026)
by: Liu, Jiaxuan, et al.
Published: (2026)
DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos
by: Li, Can, et al.
Published: (2026)
by: Li, Can, et al.
Published: (2026)
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
by: Li, Lei-lei, et al.
Published: (2025)
by: Li, Lei-lei, et al.
Published: (2025)
Similar Items
-
SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality
by: Lei, Chenyang, et al.
Published: (2024) -
SimMAT: Exploring Transferability from Vision Foundation Models to Any Image Modality
by: Lei, Chenyang, et al.
Published: (2024) -
Generative Active Learning for Image Synthesis Personalization
by: Zhang, Xulu, et al.
Published: (2024) -
Compositional Inversion for Stable Diffusion Models
by: Zhang, Xulu, et al.
Published: (2023) -
General Geometry-aware Weakly Supervised 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2024)