Saved in:
| Main Authors: | Liu, Xiangrui, Li, Haoxiang, Yang, Yezhou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.23167 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
Incorporating dense metric depth into neural 3D representations for view synthesis and relighting
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024)
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024)
X-Part: high fidelity and structure coherent shape decomposition
by: Yan, Xinhao, et al.
Published: (2025)
by: Yan, Xinhao, et al.
Published: (2025)
Advancing high-fidelity 3D and Texture Generation with 2.5D latents
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
HiFi-123: Towards High-fidelity One Image to 3D Content Generation
by: Yu, Wangbo, et al.
Published: (2023)
by: Yu, Wangbo, et al.
Published: (2023)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
by: Feinglass, Joshua, et al.
Published: (2024)
by: Feinglass, Joshua, et al.
Published: (2024)
R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model
by: Kim, Changhoon, et al.
Published: (2024)
by: Kim, Changhoon, et al.
Published: (2024)
A training-free framework for high-fidelity appearance transfer via diffusion transformers
by: Gu, Shengrong, et al.
Published: (2026)
by: Gu, Shengrong, et al.
Published: (2026)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
by: Malaviya, Vatsal, et al.
Published: (2025)
by: Malaviya, Vatsal, et al.
Published: (2025)
GHOST 2.0: generative high-fidelity one shot transfer of heads
by: Groshev, Alexander, et al.
Published: (2025)
by: Groshev, Alexander, et al.
Published: (2025)
Fine-grained subjective visual quality assessment for high-fidelity compressed images
by: Testolina, Michela, et al.
Published: (2024)
by: Testolina, Michela, et al.
Published: (2024)
EiHi Net: Out-of-Distribution Generalization Paradigm
by: Wei, Qinglai, et al.
Published: (2022)
by: Wei, Qinglai, et al.
Published: (2022)
HiNeuS: High-fidelity Neural Surface Mitigating Low-texture and Reflective Ambiguity
by: Wang, Yida, et al.
Published: (2025)
by: Wang, Yida, et al.
Published: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
by: Peirone, Simone Alberto, et al.
Published: (2025)
by: Peirone, Simone Alberto, et al.
Published: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
by: Cheng, Sheng, et al.
Published: (2024)
by: Cheng, Sheng, et al.
Published: (2024)
Virtual-reality based patient-specific simulation of spine surgical procedures: A fast, highly automated and high-fidelity system for surgical education and planning
by: Ranabhat, Raj Kumar, et al.
Published: (2026)
by: Ranabhat, Raj Kumar, et al.
Published: (2026)
RealSynCol: a high-fidelity synthetic colon dataset for 3D reconstruction applications
by: Lena, Chiara, et al.
Published: (2026)
by: Lena, Chiara, et al.
Published: (2026)
VOCAL: Visual Odometry via ContrAstive Learning
by: Huang, Chi-Yao, et al.
Published: (2025)
by: Huang, Chi-Yao, et al.
Published: (2025)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
by: Patel, Maitreya, et al.
Published: (2023)
by: Patel, Maitreya, et al.
Published: (2023)
SKoPe3D: A Synthetic Dataset for Vehicle Keypoint Perception in 3D from Traffic Monitoring Cameras
by: Pahadia, Himanshu, et al.
Published: (2023)
by: Pahadia, Himanshu, et al.
Published: (2023)
Scan Clusters, Not Pixels: A Cluster-Centric Paradigm for Efficient Ultra-high-definition Image Restoration
by: Wu, Chen, et al.
Published: (2026)
by: Wu, Chen, et al.
Published: (2026)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
ShadeBench: A Benchmark Dataset for Building Shade Simulation in Sustainable Society
by: Da, Longchao, et al.
Published: (2026)
by: Da, Longchao, et al.
Published: (2026)
Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging
by: Ye, Chongjie, et al.
Published: (2025)
by: Ye, Chongjie, et al.
Published: (2025)
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
HiLLIE: Human-in-the-Loop Training for Low-Light Image Enhancement
by: Zhao, Xiaorui, et al.
Published: (2025)
by: Zhao, Xiaorui, et al.
Published: (2025)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
by: Pathiraja, Bimsara, et al.
Published: (2025)
by: Pathiraja, Bimsara, et al.
Published: (2025)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Asynchronous Remote Sensing Time-Series Fusion for Cloud Removal and Anytime Reconstruction
by: Fallah, Forouzan, et al.
Published: (2026)
by: Fallah, Forouzan, et al.
Published: (2026)
RareFlow: Physics-Aware Flow-Matching for Cross-Sensor Super-Resolution of Rare-Earth Features
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement
by: Zhou, Kun, et al.
Published: (2024)
by: Zhou, Kun, et al.
Published: (2024)
Projected Representation Conditioning for High-fidelity Novel View Synthesis
by: Kwak, Min-Seop, et al.
Published: (2026)
by: Kwak, Min-Seop, et al.
Published: (2026)
Robust soybean seed yield estimation using high-throughput ground robot videos
by: Feng, Jiale, et al.
Published: (2024)
by: Feng, Jiale, et al.
Published: (2024)
Recent Event Camera Innovations: A Survey
by: Chakravarthi, Bharatesh, et al.
Published: (2024)
by: Chakravarthi, Bharatesh, et al.
Published: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
MentalBlackboard: Evaluating Spatial Visualization via Mathematical Transformations
by: Yilmaz, Nilay, et al.
Published: (2026)
by: Yilmaz, Nilay, et al.
Published: (2026)
Similar Items
-
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
by: Liu, Xiangrui, et al.
Published: (2025) -
Incorporating dense metric depth into neural 3D representations for view synthesis and relighting
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024) -
X-Part: high fidelity and structure coherent shape decomposition
by: Yan, Xinhao, et al.
Published: (2025) -
Advancing high-fidelity 3D and Texture Generation with 2.5D latents
by: Yang, Xin, et al.
Published: (2025) -
HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
by: Xu, Yifang, et al.
Published: (2025)