Understanding Generative AI Capabilities in Everyday Image Editing Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Taesiri, Mohammad Reza, Collins, Brandon, Bolton, Logan, Lai, Viet Dac, Dernoncourt, Franck, Bui, Trung, Nguyen, Anh Totti |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026)
by: Collins, Brandon, et al.
Published: (2026)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
anguyen8/vision-llms-are-blind: official
by: Pooyan R, et al.
Published: (2026)
by: Pooyan R, et al.
Published: (2026)
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
SlimLM: An Efficient Small Language Model for On-Device Document Assistance
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
Vision Language Models are Biased
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models
by: Man, Hieu, et al.
Published: (2025)
by: Man, Hieu, et al.
Published: (2025)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
CORG: Generating Answers from Complex, Interrelated Contexts
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
An Analysis of Multilingual FActScore
by: Vu, Kim Trong, et al.
Published: (2024)
by: Vu, Kim Trong, et al.
Published: (2024)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
PageGuide: Browser extension to assist users in navigating a webpage and locating information
by: Nguyen, Tin, et al.
Published: (2026)
by: Nguyen, Tin, et al.
Published: (2026)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
DynaSaur: Large Language Agents Beyond Predefined Actions
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
Task-driven Layerwise Additive Activation Intervention
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
Explaining Graph Neural Networks via Structure-aware Interaction Index
by: Bui, Ngoc, et al.
Published: (2024)
by: Bui, Ngoc, et al.
Published: (2024)
Agentic Planning with Reasoning for Image Styling via Offline RL
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
Lizard: An Efficient Linearization Framework for Large Language Models
by: Van Nguyen, Chien, et al.
Published: (2025)
by: Van Nguyen, Chien, et al.
Published: (2025)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
by: Ngo, Nghia Trung, et al.
Published: (2025)
by: Ngo, Nghia Trung, et al.
Published: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
by: Ngo, Nghia Trung, et al.
Published: (2024)
by: Ngo, Nghia Trung, et al.
Published: (2024)
Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI
by: Man, Hieu, et al.
Published: (2026)
by: Man, Hieu, et al.
Published: (2026)
ULLME: A Unified Framework for Large Language Model Embeddings with Generation-Augmented Learning
by: Man, Hieu, et al.
Published: (2024)
by: Man, Hieu, et al.
Published: (2024)
VideoGameBunny: Towards vision assistants for video games
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
Generative Conditional Distributions by Neural (Entropic) Optimal Transport
by: Nguyen, Bao, et al.
Published: (2024)
by: Nguyen, Bao, et al.
Published: (2024)
Uniqueness of tangent currents for positive closed currents
by: Nguyen, Viet-Anh, et al.
Published: (2025)
by: Nguyen, Viet-Anh, et al.
Published: (2025)
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
by: Dongre, Vardhan, et al.
Published: (2025)
by: Dongre, Vardhan, et al.
Published: (2025)
Structured Pruning for Diverse Best-of-N Reasoning Optimization
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
Leveraging Habitat Information for Fine-grained Bird Identification
by: Nguyen, Tin, et al.
Published: (2023)
by: Nguyen, Tin, et al.
Published: (2023)
A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
by: Elmoghany, Mohamed, et al.
Published: (2025)
by: Elmoghany, Mohamed, et al.
Published: (2025)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)
by: Nguyen, Viet, et al.
Published: (2024)
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
Unleashing Creativity for Sustainable Development: From Everyday Imagination to AI ‐Enabled Innovation
by: Mohammad Reza Salamat
Published: (2026)
by: Mohammad Reza Salamat
Published: (2026)
Retrieval Augmented Generation for Domain-specific Question Answering
by: Sharma, Sanat, et al.
Published: (2024)
by: Sharma, Sanat, et al.
Published: (2024)
Similar Items
-
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026) -
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025) -
anguyen8/vision-llms-are-blind: official
by: Pooyan R, et al.
Published: (2026) -
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024) -
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)