Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Ziyi, Gao, Zhi, Chen, Jin, Zhao, Qingjie, Wu, Xinxiao, Luo, Jiebo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DiffCamera: Arbitrary Refocusing on Images
por: Wang, Yiyang, et al.
Publicado: (2025)
por: Wang, Yiyang, et al.
Publicado: (2025)
Learning to Refocus with Video Diffusion Models
por: Tedla, SaiKiran, et al.
Publicado: (2025)
por: Tedla, SaiKiran, et al.
Publicado: (2025)
Robustifying Point Cloud Networks by Refocusing
por: Levi, Meir Yossef, et al.
Publicado: (2023)
por: Levi, Meir Yossef, et al.
Publicado: (2023)
LaRe: Latent Refocusing for Multimodal Reasoning
por: Ma, Jizheng, et al.
Publicado: (2025)
por: Ma, Jizheng, et al.
Publicado: (2025)
Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
por: Li, Jiahao, et al.
Publicado: (2025)
por: Li, Jiahao, et al.
Publicado: (2025)
Generative Refocusing: Flexible Defocus Control from a Single Image
por: Mu, Chun-Wei Tuan, et al.
Publicado: (2025)
por: Mu, Chun-Wei Tuan, et al.
Publicado: (2025)
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
por: Yang, Yang, et al.
Publicado: (2025)
por: Yang, Yang, et al.
Publicado: (2025)
StarVid: Enhancing Semantic Alignment in Video Diffusion Models via Spatial and SynTactic Guided Attention Refocusing
por: Li, Yuanhang, et al.
Publicado: (2024)
por: Li, Yuanhang, et al.
Publicado: (2024)
VUDG: A Dataset for Video Understanding Domain Generalization
por: Wang, Ziyi, et al.
Publicado: (2025)
por: Wang, Ziyi, et al.
Publicado: (2025)
Arbitrary Volumetric Refocusing of Dense and Sparse Light Fields
por: Samarakoon, Tharindu, et al.
Publicado: (2025)
por: Samarakoon, Tharindu, et al.
Publicado: (2025)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
por: Wang, Yongqi, et al.
Publicado: (2024)
por: Wang, Yongqi, et al.
Publicado: (2024)
Small Clips, Big Gains: Learning Long-Range Refocused Temporal Information for Video Super-Resolution
por: Zhou, Xingyu, et al.
Publicado: (2025)
por: Zhou, Xingyu, et al.
Publicado: (2025)
Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning
por: Shen, Ruolin, et al.
Publicado: (2025)
por: Shen, Ruolin, et al.
Publicado: (2025)
DOF-GS:Adjustable Depth-of-Field 3D Gaussian Splatting for Post-Capture Refocusing, Defocus Rendering and Blur Removal
por: Wang, Yujie, et al.
Publicado: (2024)
por: Wang, Yujie, et al.
Publicado: (2024)
VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
por: Qi, Yayun, et al.
Publicado: (2024)
por: Qi, Yayun, et al.
Publicado: (2024)
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
por: Wang, Yongqi, et al.
Publicado: (2025)
por: Wang, Yongqi, et al.
Publicado: (2025)
LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
por: Shang, Zirui, et al.
Publicado: (2025)
por: Shang, Zirui, et al.
Publicado: (2025)
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
por: Tian, Mengxiao, et al.
Publicado: (2025)
por: Tian, Mengxiao, et al.
Publicado: (2025)
On Inductive Biases That Enable Generalization of Diffusion Transformers
por: An, Jie, et al.
Publicado: (2024)
por: An, Jie, et al.
Publicado: (2024)
Self-Ensemble Post Learning for Noisy Domain Generalization
por: Lu, Wang, et al.
Publicado: (2025)
por: Lu, Wang, et al.
Publicado: (2025)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
por: Zhang, Bofei, et al.
Publicado: (2025)
por: Zhang, Bofei, et al.
Publicado: (2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
por: Zhou, Zhenghong, et al.
Publicado: (2024)
por: Zhou, Zhenghong, et al.
Publicado: (2024)
Vision Transformers are Circulant Attention Learners
por: Han, Dongchen, et al.
Publicado: (2025)
por: Han, Dongchen, et al.
Publicado: (2025)
Pixel-Level Domain Adaptation: A New Perspective for Enhancing Weakly Supervised Semantic Segmentation
por: Du, Ye, et al.
Publicado: (2024)
por: Du, Ye, et al.
Publicado: (2024)
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
por: Yuan, Shenghai, et al.
Publicado: (2024)
por: Yuan, Shenghai, et al.
Publicado: (2024)
De-Simplifying Pseudo Labels to Enhancing Domain Adaptive Object Detection
por: Fu, Zehua, et al.
Publicado: (2025)
por: Fu, Zehua, et al.
Publicado: (2025)
DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object Detection
por: Feng, Yongchao, et al.
Publicado: (2023)
por: Feng, Yongchao, et al.
Publicado: (2023)
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
por: Pan, Jiadong, et al.
Publicado: (2024)
por: Pan, Jiadong, et al.
Publicado: (2024)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
por: Zeng, Ziyun, et al.
Publicado: (2025)
por: Zeng, Ziyun, et al.
Publicado: (2025)
PACF: Prototype Augmented Compact Features for Improving Domain Adaptive Object Detection
por: Liu, Chenguang, et al.
Publicado: (2025)
por: Liu, Chenguang, et al.
Publicado: (2025)
Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
por: Yu, Yongsheng, et al.
Publicado: (2024)
por: Yu, Yongsheng, et al.
Publicado: (2024)
Adaptive Model Ensemble for Continual Learning
por: Mao, Yuchuan, et al.
Publicado: (2025)
por: Mao, Yuchuan, et al.
Publicado: (2025)
A Versatile Multimodal Agent for Multimedia Content Generation
por: Zhang, Daoan, et al.
Publicado: (2026)
por: Zhang, Daoan, et al.
Publicado: (2026)
Semantic Enhanced Few-shot Object Detection
por: Wang, Zheng, et al.
Publicado: (2024)
por: Wang, Zheng, et al.
Publicado: (2024)
DiffCLIP: Leveraging Stable Diffusion for Language Grounded 3D Classification
por: Shen, Sitian, et al.
Publicado: (2023)
por: Shen, Sitian, et al.
Publicado: (2023)
Video Summarization using Denoising Diffusion Probabilistic Model
por: Shang, Zirui, et al.
Publicado: (2024)
por: Shang, Zirui, et al.
Publicado: (2024)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
por: Lu, Beijia, et al.
Publicado: (2025)
por: Lu, Beijia, et al.
Publicado: (2025)
Data-free Multi-label Image Recognition via LLM-powered Prompt Tuning
por: Yang, Shuo, et al.
Publicado: (2024)
por: Yang, Shuo, et al.
Publicado: (2024)
EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
por: Yang, Bingyu, et al.
Publicado: (2025)
por: Yang, Bingyu, et al.
Publicado: (2025)
Ejemplares similares
-
DiffCamera: Arbitrary Refocusing on Images
por: Wang, Yiyang, et al.
Publicado: (2025) -
Learning to Refocus with Video Diffusion Models
por: Tedla, SaiKiran, et al.
Publicado: (2025) -
Robustifying Point Cloud Networks by Refocusing
por: Levi, Meir Yossef, et al.
Publicado: (2023) -
LaRe: Latent Refocusing for Multimodal Reasoning
por: Ma, Jizheng, et al.
Publicado: (2025) -
Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
por: Li, Jiahao, et al.
Publicado: (2025)