Rethinking FID Through the Geometry of the Reference Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yunghee, Pak, Byeonghyun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
Tortoise and Hare Guidance: Accelerating Diffusion Model Inference with Multirate Integration
by: Lee, Yunghee, et al.
Published: (2025)
by: Lee, Yunghee, et al.
Published: (2025)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
Seeing Through the Mask: Rethinking Adversarial Examples for CAPTCHAs
by: Jabary, Yahya, et al.
Published: (2024)
by: Jabary, Yahya, et al.
Published: (2024)
Rethinking Visual Counterfactual Explanations Through Region Constraint
by: Sobieski, Bartlomiej, et al.
Published: (2024)
by: Sobieski, Bartlomiej, et al.
Published: (2024)
Rethinking RGB-D Fusion for Semantic Segmentation in Surgical Datasets
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
GeoDM: Geometry-aware Distribution Matching for Dataset Distillation
by: Li, Xuhui, et al.
Published: (2025)
by: Li, Xuhui, et al.
Published: (2025)
DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
Exploring Spatial Language Grounding Through Referring Expressions
by: Tumu, Akshar, et al.
Published: (2025)
by: Tumu, Akshar, et al.
Published: (2025)
Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting Code
by: Lin, Haobo, et al.
Published: (2026)
by: Lin, Haobo, et al.
Published: (2026)
SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
by: Lee, Jungho, et al.
Published: (2025)
by: Lee, Jungho, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
by: Pak, Byeonghyun, et al.
Published: (2024)
by: Pak, Byeonghyun, et al.
Published: (2024)
Open Set Recognition for Endoscopic Image Classification: A Deep Learning Approach on the Kvasir Dataset
by: Moazzami, Kasra, et al.
Published: (2025)
by: Moazzami, Kasra, et al.
Published: (2025)
Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
by: Jeon, Wooseok, et al.
Published: (2026)
by: Jeon, Wooseok, et al.
Published: (2026)
ReferGPT: Towards Zero-Shot Referring Multi-Object Tracking
by: Chamiti, Tzoulio, et al.
Published: (2025)
by: Chamiti, Tzoulio, et al.
Published: (2025)
Rethinking Genomic Modeling Through Optical Character Recognition
by: Xiang, Hongxin, et al.
Published: (2026)
by: Xiang, Hongxin, et al.
Published: (2026)
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
by: Jiang, Le, et al.
Published: (2026)
by: Jiang, Le, et al.
Published: (2026)
Self-training Room Layout Estimation via Geometry-aware Ray-casting
by: Solarte, Bolivar, et al.
Published: (2024)
by: Solarte, Bolivar, et al.
Published: (2024)
Rethink MAE with Linear Time-Invariant Dynamics
by: Wang, Zice
Published: (2026)
by: Wang, Zice
Published: (2026)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
Beyond the Failures: Rethinking Foundation Models in Pathology
by: Tizhoosh, Hamid R.
Published: (2025)
by: Tizhoosh, Hamid R.
Published: (2025)
Rethinking Unsupervised Domain Adaptation for Semantic Segmentation
by: Wang, Zhijie, et al.
Published: (2022)
by: Wang, Zhijie, et al.
Published: (2022)
Rethinking Alignment and Uniformity in Unsupervised Semantic Segmentation
by: Zhang, Daoan, et al.
Published: (2022)
by: Zhang, Daoan, et al.
Published: (2022)
Rethinking the Sample Relations for Few-Shot Classification
by: Yin, Guowei, et al.
Published: (2025)
by: Yin, Guowei, et al.
Published: (2025)
Rethinking Visual Information Processing in Multimodal LLMs
by: Kim, Dongwan, et al.
Published: (2025)
by: Kim, Dongwan, et al.
Published: (2025)
VACoT: Rethinking Visual Data Augmentation with VLMs
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
by: Wei, Meng, et al.
Published: (2026)
by: Wei, Meng, et al.
Published: (2026)
Bridging the Domain Gap: A Simple Domain Matching Method for Reference-based Image Super-Resolution in Remote Sensing
by: Min, Jeongho, et al.
Published: (2024)
by: Min, Jeongho, et al.
Published: (2024)
CLOFAI: A Dataset of Real And Fake Image Classification Tasks for Continual Learning
by: Doherty, William, et al.
Published: (2025)
by: Doherty, William, et al.
Published: (2025)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
SF-Mamba: Rethinking State Space Model for Vision
by: Yoshimura, Masakazu, et al.
Published: (2026)
by: Yoshimura, Masakazu, et al.
Published: (2026)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026)
by: Feng, Zhe, et al.
Published: (2026)
Rethinking Cross-Layer Information Routing in Diffusion Transformers
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
Rethinking Token Reduction for Large Vision-Language Models
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
by: Shen, Dazhong, et al.
Published: (2024)
by: Shen, Dazhong, et al.
Published: (2024)
Rethinking Causal Mask Attention for Vision-Language Inference
by: Pei, Xiaohuan, et al.
Published: (2025)
by: Pei, Xiaohuan, et al.
Published: (2025)
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On
by: Na, Kihyun, et al.
Published: (2025)
by: Na, Kihyun, et al.
Published: (2025)
Similar Items
-
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026) -
Tortoise and Hare Guidance: Accelerating Diffusion Model Inference with Multirate Integration
by: Lee, Yunghee, et al.
Published: (2025) -
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026) -
Seeing Through the Mask: Rethinking Adversarial Examples for CAPTCHAs
by: Jabary, Yahya, et al.
Published: (2024) -
Rethinking Visual Counterfactual Explanations Through Region Constraint
by: Sobieski, Bartlomiej, et al.
Published: (2024)