MAR-MAER: Metric-Aware and Ambiguity-Adaptive Autoregressive Image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Dong, Kai, Bai, Tingting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Context-Aware Full Body Anonymization using Text-to-Image Diffusion Models
por: Zwick, Pascal, et al.
Publicado: (2024)
por: Zwick, Pascal, et al.
Publicado: (2024)
Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation
por: Du, Shuangli, et al.
Publicado: (2025)
por: Du, Shuangli, et al.
Publicado: (2025)
Generating Image Adversarial Examples by Embedding Digital Watermarks
por: Xiang, Yuexin, et al.
Publicado: (2020)
por: Xiang, Yuexin, et al.
Publicado: (2020)
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
por: Ren, Yumeng, et al.
Publicado: (2025)
por: Ren, Yumeng, et al.
Publicado: (2025)
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
por: Liu, Ying, et al.
Publicado: (2025)
por: Liu, Ying, et al.
Publicado: (2025)
EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts
por: Mukherjee, Kushin, et al.
Publicado: (2025)
por: Mukherjee, Kushin, et al.
Publicado: (2025)
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
por: Mathys, Melanie, et al.
Publicado: (2024)
por: Mathys, Melanie, et al.
Publicado: (2024)
Measuring Diversity in Co-creative Image Generation
por: Ibarrola, Francisco, et al.
Publicado: (2024)
por: Ibarrola, Francisco, et al.
Publicado: (2024)
Synthetic Image Generation in Cyber Influence Operations: An Emergent Threat?
por: Mathys, Melanie, et al.
Publicado: (2024)
por: Mathys, Melanie, et al.
Publicado: (2024)
Exploring Transfer Learning for Deep Learning Polyp Detection in Colonoscopy Images Using YOLOv8
por: Vazquez, Fabian, et al.
Publicado: (2025)
por: Vazquez, Fabian, et al.
Publicado: (2025)
CerberusDet: Unified Multi-Dataset Object Detection
por: Tolstykh, Irina, et al.
Publicado: (2024)
por: Tolstykh, Irina, et al.
Publicado: (2024)
Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models
por: Mandelli, Lorenzo, et al.
Publicado: (2024)
por: Mandelli, Lorenzo, et al.
Publicado: (2024)
Generative inpainting of incomplete Euclidean distance matrices of trajectories generated by a fractional Brownian motion
por: Lobashev, Alexander, et al.
Publicado: (2024)
por: Lobashev, Alexander, et al.
Publicado: (2024)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
por: He, Mengqi, et al.
Publicado: (2025)
por: He, Mengqi, et al.
Publicado: (2025)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
por: Zhang, Peng-Fei, et al.
Publicado: (2024)
por: Zhang, Peng-Fei, et al.
Publicado: (2024)
Measuring proximity to standard planes during fetal brain ultrasound scanning
por: Di Vece, Chiara, et al.
Publicado: (2024)
por: Di Vece, Chiara, et al.
Publicado: (2024)
Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
por: Lu, Zhuqiang, et al.
Publicado: (2023)
por: Lu, Zhuqiang, et al.
Publicado: (2023)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
por: Altamirano-Gomez, Gerardo, et al.
Publicado: (2023)
por: Altamirano-Gomez, Gerardo, et al.
Publicado: (2023)
SeNeDiF-OOD: Semantic Nested Dichotomy Fusion for Out-of-Distribution Detection Methodology in Open-World Classification. A Case Study on Monument Style Classification
por: Antequera-Sánchez, Ignacio, et al.
Publicado: (2026)
por: Antequera-Sánchez, Ignacio, et al.
Publicado: (2026)
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
por: Yan, Wen, et al.
Publicado: (2026)
por: Yan, Wen, et al.
Publicado: (2026)
CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
por: Li, Jiaxuan, et al.
Publicado: (2025)
por: Li, Jiaxuan, et al.
Publicado: (2025)
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
por: Tang, MingZe, et al.
Publicado: (2025)
por: Tang, MingZe, et al.
Publicado: (2025)
Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
por: Nguyen, Viet Dung, et al.
Publicado: (2026)
por: Nguyen, Viet Dung, et al.
Publicado: (2026)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
por: Tangtartharakul, Gene, et al.
Publicado: (2026)
por: Tangtartharakul, Gene, et al.
Publicado: (2026)
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
por: Liu, Liangjin, et al.
Publicado: (2025)
por: Liu, Liangjin, et al.
Publicado: (2025)
A Review of Pseudo-Labeling for Computer Vision
por: Kage, Patrick, et al.
Publicado: (2024)
por: Kage, Patrick, et al.
Publicado: (2024)
Sparse vs Contiguous Adversarial Pixel Perturbations in Multimodal Models: An Empirical Analysis
por: Botocan, Cristian-Alexandru, et al.
Publicado: (2024)
por: Botocan, Cristian-Alexandru, et al.
Publicado: (2024)
Towards Interpretable Visual Decoding with Attention to Brain Representations
por: Feng, Pinyuan, et al.
Publicado: (2025)
por: Feng, Pinyuan, et al.
Publicado: (2025)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
por: Wang, Han, et al.
Publicado: (2024)
por: Wang, Han, et al.
Publicado: (2024)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
por: Ma, Chong, et al.
Publicado: (2024)
por: Ma, Chong, et al.
Publicado: (2024)
Lookism: The overlooked bias in computer vision
por: Gulati, Aditya, et al.
Publicado: (2024)
por: Gulati, Aditya, et al.
Publicado: (2024)
FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare
por: Lekadir, Karim, et al.
Publicado: (2023)
por: Lekadir, Karim, et al.
Publicado: (2023)
Parameterizing Dataset Distillation via Gaussian Splatting
por: Jiang, Chenyang, et al.
Publicado: (2025)
por: Jiang, Chenyang, et al.
Publicado: (2025)
Deep Domain Adaptation: A Sim2Real Neural Approach for Improving Eye-Tracking Systems
por: Nguyen, Viet Dung, et al.
Publicado: (2024)
por: Nguyen, Viet Dung, et al.
Publicado: (2024)
Convolutional Neural Networks Can (Meta-)Learn the Same-Different Relation
por: Gupta, Max, et al.
Publicado: (2025)
por: Gupta, Max, et al.
Publicado: (2025)
Visual Language Models show widespread visual deficits on neuropsychological tests
por: Tangtartharakul, Gene, et al.
Publicado: (2025)
por: Tangtartharakul, Gene, et al.
Publicado: (2025)
Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation
por: Kuprashevich, Maksim, et al.
Publicado: (2024)
por: Kuprashevich, Maksim, et al.
Publicado: (2024)
VACoDe: Visual Augmented Contrastive Decoding
por: Kim, Sihyeon, et al.
Publicado: (2024)
por: Kim, Sihyeon, et al.
Publicado: (2024)
Towards Infusing Auxiliary Knowledge for Distracted Driver Detection
por: Balappanawar, Ishwar B, et al.
Publicado: (2024)
por: Balappanawar, Ishwar B, et al.
Publicado: (2024)
Scene-wise Adaptive Network for Dynamic Cold-start Scenes Optimization in CTR Prediction
por: Li, Wenhao, et al.
Publicado: (2024)
por: Li, Wenhao, et al.
Publicado: (2024)
Ejemplares similares
-
Context-Aware Full Body Anonymization using Text-to-Image Diffusion Models
por: Zwick, Pascal, et al.
Publicado: (2024) -
Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation
por: Du, Shuangli, et al.
Publicado: (2025) -
Generating Image Adversarial Examples by Embedding Digital Watermarks
por: Xiang, Yuexin, et al.
Publicado: (2020) -
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
por: Ren, Yumeng, et al.
Publicado: (2025) -
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
por: Liu, Ying, et al.
Publicado: (2025)