Diffusion-based Visual Anagram as Multi-task Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Zhiyuan, Chen, Yinhe, Gao, Huan-ang, Zhao, Weiyan, Zhang, Guiyu, Zhao, Hao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FairDiff: Fair Segmentation with Point-Image Diffusion
por: Li, Wenyi, et al.
Publicado: (2024)
por: Li, Wenyi, et al.
Publicado: (2024)
Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling
por: Zhang, Guiyu, et al.
Publicado: (2024)
por: Zhang, Guiyu, et al.
Publicado: (2024)
Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models
por: Geng, Daniel, et al.
Publicado: (2023)
por: Geng, Daniel, et al.
Publicado: (2023)
Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence Loss
por: Zhang, Yifei, et al.
Publicado: (2024)
por: Zhang, Yifei, et al.
Publicado: (2024)
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
por: Gao, Huan-ang, et al.
Publicado: (2024)
por: Gao, Huan-ang, et al.
Publicado: (2024)
Training-Free Model Merging for Multi-target Domain Adaptation
por: Li, Wenyi, et al.
Publicado: (2024)
por: Li, Wenyi, et al.
Publicado: (2024)
Challenger: Affordable Adversarial Driving Video Generation
por: Xu, Zhiyuan, et al.
Publicado: (2025)
por: Xu, Zhiyuan, et al.
Publicado: (2025)
Alias-free 4D Gaussian Splatting
por: Chen, Zilong, et al.
Publicado: (2025)
por: Chen, Zilong, et al.
Publicado: (2025)
AVD2: Accident Video Diffusion for Accident Video Description
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
por: Doshi, Fenil R., et al.
Publicado: (2025)
por: Doshi, Fenil R., et al.
Publicado: (2025)
PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model
por: Gao, Mingju, et al.
Publicado: (2025)
por: Gao, Mingju, et al.
Publicado: (2025)
FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
por: Li, Jinwei, et al.
Publicado: (2025)
por: Li, Jinwei, et al.
Publicado: (2025)
SGOR: Outlier Removal by Leveraging Semantic and Geometric Information for Robust Point Cloud Registration
por: Zhao, Guiyu, et al.
Publicado: (2024)
por: Zhao, Guiyu, et al.
Publicado: (2024)
P-MapNet: Far-seeing Map Generator Enhanced by both SDMap and HDMap Priors
por: Jiang, Zhou, et al.
Publicado: (2024)
por: Jiang, Zhou, et al.
Publicado: (2024)
SA-GS: Scale-Adaptive Gaussian Splatting for Training-Free Anti-Aliasing
por: Song, Xiaowei, et al.
Publicado: (2024)
por: Song, Xiaowei, et al.
Publicado: (2024)
Progressive Correspondence Regenerator for Robust 3D Registration
por: Zhao, Guiyu, et al.
Publicado: (2025)
por: Zhao, Guiyu, et al.
Publicado: (2025)
RGM: Reconstructing High-fidelity 3D Car Assets with Relightable 3D-GS Generative Model from a Single Image
por: Chen, Xiaoxue, et al.
Publicado: (2024)
por: Chen, Xiaoxue, et al.
Publicado: (2024)
Benchmarking PhD-Level Coding in 3D Geometric Computer Vision
por: Li, Wenyi, et al.
Publicado: (2026)
por: Li, Wenyi, et al.
Publicado: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
VRHCF: Cross-Source Point Cloud Registration via Voxel Representation and Hierarchical Correspondence Filtering
por: Zhao, Guiyu, et al.
Publicado: (2024)
por: Zhao, Guiyu, et al.
Publicado: (2024)
Cross-PCR: A Robust Cross-Source Point Cloud Registration Framework
por: Zhao, Guiyu, et al.
Publicado: (2024)
por: Zhao, Guiyu, et al.
Publicado: (2024)
Rip-NeRF: Anti-aliasing Radiance Fields with Ripmap-Encoded Platonic Solids
por: Liu, Junchen, et al.
Publicado: (2024)
por: Liu, Junchen, et al.
Publicado: (2024)
Elite360M: Efficient 360 Multi-task Learning via Bi-projection Fusion and Cross-task Collaboration
por: Ai, Hao, et al.
Publicado: (2024)
por: Ai, Hao, et al.
Publicado: (2024)
Latency-aware Road Anomaly Segmentation in Videos: A Photorealistic Dataset and New Metrics
por: Tian, Beiwen, et al.
Publicado: (2024)
por: Tian, Beiwen, et al.
Publicado: (2024)
Prototype-Based Low Altitude UAV Semantic Segmentation
por: Zhang, Da, et al.
Publicado: (2026)
por: Zhang, Da, et al.
Publicado: (2026)
Cross-Layer Feature Pyramid Transformer for Small Object Detection in Aerial Images
por: Du, Zewen, et al.
Publicado: (2024)
por: Du, Zewen, et al.
Publicado: (2024)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
por: Ding, Kairui, et al.
Publicado: (2024)
por: Ding, Kairui, et al.
Publicado: (2024)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
por: Zhao, Canyu, et al.
Publicado: (2025)
por: Zhao, Canyu, et al.
Publicado: (2025)
Learning an Implicit Physics Model for Image-based Fluid Simulation
por: Jia, Emily Yue-Ting, et al.
Publicado: (2025)
por: Jia, Emily Yue-Ting, et al.
Publicado: (2025)
OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
por: Guo, Qi, et al.
Publicado: (2026)
por: Guo, Qi, et al.
Publicado: (2026)
Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail
por: Chen, Mingjin, et al.
Publicado: (2024)
por: Chen, Mingjin, et al.
Publicado: (2024)
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
por: Gao, Mingju, et al.
Publicado: (2026)
por: Gao, Mingju, et al.
Publicado: (2026)
MultiTaskVIF: Segmentation-oriented visible and infrared image fusion via multi-task learning
por: Zhao, Zixian, et al.
Publicado: (2025)
por: Zhao, Zixian, et al.
Publicado: (2025)
Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting
por: Hsu, Tsuheng, et al.
Publicado: (2026)
por: Hsu, Tsuheng, et al.
Publicado: (2026)
Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation
por: Pan, Jiadong, et al.
Publicado: (2025)
por: Pan, Jiadong, et al.
Publicado: (2025)
Reusing Attention for One-stage Lane Topology Understanding
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment
por: Wang, Bin, et al.
Publicado: (2025)
por: Wang, Bin, et al.
Publicado: (2025)
Delving into Mapping Uncertainty for Mapless Trajectory Prediction
por: Zhang, Zongzheng, et al.
Publicado: (2025)
por: Zhang, Zongzheng, et al.
Publicado: (2025)
One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models
por: Zhao, Jiale, et al.
Publicado: (2025)
por: Zhao, Jiale, et al.
Publicado: (2025)
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
por: Wang, Xinyi, et al.
Publicado: (2025)
por: Wang, Xinyi, et al.
Publicado: (2025)
Ejemplares similares
-
FairDiff: Fair Segmentation with Point-Image Diffusion
por: Li, Wenyi, et al.
Publicado: (2024) -
Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling
por: Zhang, Guiyu, et al.
Publicado: (2024) -
Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models
por: Geng, Daniel, et al.
Publicado: (2023) -
Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence Loss
por: Zhang, Yifei, et al.
Publicado: (2024) -
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
por: Gao, Huan-ang, et al.
Publicado: (2024)