InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Jinbin, He, Wenbin, Gou, Liang, Ren, Liu, Bryan, Chris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InFiConD: Interactive No-code Fine-tuning with Concept-based Knowledge Distillation
di: Huang, Jinbin, et al.
Pubblicazione: (2024)
di: Huang, Jinbin, et al.
Pubblicazione: (2024)
ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale
di: Huang, Jinbin, et al.
Pubblicazione: (2024)
di: Huang, Jinbin, et al.
Pubblicazione: (2024)
AttributionScanner: A Visual Analytics System for Model Validation with Metadata-Free Slice Finding
di: Xuan, Xiwei, et al.
Pubblicazione: (2024)
di: Xuan, Xiwei, et al.
Pubblicazione: (2024)
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
di: Panigrahi, Indu, et al.
Pubblicazione: (2025)
di: Panigrahi, Indu, et al.
Pubblicazione: (2025)
InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
di: Lin, Yukang, et al.
Pubblicazione: (2025)
di: Lin, Yukang, et al.
Pubblicazione: (2025)
SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision
di: Lesner, Jasmine, et al.
Pubblicazione: (2026)
di: Lesner, Jasmine, et al.
Pubblicazione: (2026)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
di: Li, Chentao, et al.
Pubblicazione: (2026)
di: Li, Chentao, et al.
Pubblicazione: (2026)
VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
di: Yan, Xinyuan, et al.
Pubblicazione: (2025)
di: Yan, Xinyuan, et al.
Pubblicazione: (2025)
Vision Language Models as Values Detectors
di: Abbo, Giulio Antonio, et al.
Pubblicazione: (2025)
di: Abbo, Giulio Antonio, et al.
Pubblicazione: (2025)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
di: Huang, Yifei, et al.
Pubblicazione: (2025)
di: Huang, Yifei, et al.
Pubblicazione: (2025)
Do Vision Language Models Understand Human Engagement in Games?
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
di: Verma, Arnav, et al.
Pubblicazione: (2025)
di: Verma, Arnav, et al.
Pubblicazione: (2025)
Panda or not Panda? Understanding Adversarial Attacks with Interactive Visualization
di: You, Yuzhe, et al.
Pubblicazione: (2023)
di: You, Yuzhe, et al.
Pubblicazione: (2023)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
Acoustic Field Video for Multimodal Scene Understanding
di: Kim, Daehwa, et al.
Pubblicazione: (2026)
di: Kim, Daehwa, et al.
Pubblicazione: (2026)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
di: Nowicki, Filip, et al.
Pubblicazione: (2026)
di: Nowicki, Filip, et al.
Pubblicazione: (2026)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
di: Zheng, Ruizhe, et al.
Pubblicazione: (2024)
di: Zheng, Ruizhe, et al.
Pubblicazione: (2024)
HarassGuard: Detecting Harassment Behaviors in Social Virtual Reality with Vision-Language Models
di: Lee, Junhee, et al.
Pubblicazione: (2026)
di: Lee, Junhee, et al.
Pubblicazione: (2026)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
di: Fei, Hao, et al.
Pubblicazione: (2024)
di: Fei, Hao, et al.
Pubblicazione: (2024)
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference
di: Zhu, Yitong, et al.
Pubblicazione: (2025)
di: Zhu, Yitong, et al.
Pubblicazione: (2025)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
di: Duan, Lin, et al.
Pubblicazione: (2025)
di: Duan, Lin, et al.
Pubblicazione: (2025)
ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality
di: Khan, Dawar, et al.
Pubblicazione: (2026)
di: Khan, Dawar, et al.
Pubblicazione: (2026)
UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
di: Han, Tianhao, et al.
Pubblicazione: (2026)
di: Han, Tianhao, et al.
Pubblicazione: (2026)
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
di: Li, Xiao, et al.
Pubblicazione: (2025)
di: Li, Xiao, et al.
Pubblicazione: (2025)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
di: Garg, Kapil, et al.
Pubblicazione: (2025)
di: Garg, Kapil, et al.
Pubblicazione: (2025)
A Convolution-Based Gait Asymmetry Metric for Inter-Limb Synergistic Coordination
di: Fukino, Go, et al.
Pubblicazione: (2025)
di: Fukino, Go, et al.
Pubblicazione: (2025)
Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
di: Wu, Runtong, et al.
Pubblicazione: (2025)
di: Wu, Runtong, et al.
Pubblicazione: (2025)
Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models
di: Gallardo, Rodrigo, et al.
Pubblicazione: (2025)
di: Gallardo, Rodrigo, et al.
Pubblicazione: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
VFA: Vision Frequency Analysis of Foundation Models and Human
di: Darvishi-Bayazi, Mohammad-Javad, et al.
Pubblicazione: (2024)
di: Darvishi-Bayazi, Mohammad-Javad, et al.
Pubblicazione: (2024)
3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications
di: Corral-Soto, Eduardo R., et al.
Pubblicazione: (2024)
di: Corral-Soto, Eduardo R., et al.
Pubblicazione: (2024)
Weak-Annotation of HAR Datasets using Vision Foundation Models
di: Bock, Marius, et al.
Pubblicazione: (2024)
di: Bock, Marius, et al.
Pubblicazione: (2024)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
di: Foteinopoulou, Niki Maria, et al.
Pubblicazione: (2023)
di: Foteinopoulou, Niki Maria, et al.
Pubblicazione: (2023)
Achieving Effective Virtual Reality Interactions via Acoustic Gesture Recognition based on Large Language Models
di: Zhang, Xijie, et al.
Pubblicazione: (2025)
di: Zhang, Xijie, et al.
Pubblicazione: (2025)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
di: Niu, Runliang, et al.
Pubblicazione: (2024)
di: Niu, Runliang, et al.
Pubblicazione: (2024)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2025)
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2025)
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
di: Wang, Yuxiao, et al.
Pubblicazione: (2025)
di: Wang, Yuxiao, et al.
Pubblicazione: (2025)
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
di: Lin, Bo, et al.
Pubblicazione: (2024)
di: Lin, Bo, et al.
Pubblicazione: (2024)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
di: Li, Zisu, et al.
Pubblicazione: (2025)
di: Li, Zisu, et al.
Pubblicazione: (2025)
Viewpoint Recommendation for Point Cloud Labeling through Interaction Cost Modeling
di: Zhang, Yu, et al.
Pubblicazione: (2026)
di: Zhang, Yu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
InFiConD: Interactive No-code Fine-tuning with Concept-based Knowledge Distillation
di: Huang, Jinbin, et al.
Pubblicazione: (2024) -
ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale
di: Huang, Jinbin, et al.
Pubblicazione: (2024) -
AttributionScanner: A Visual Analytics System for Model Validation with Metadata-Free Slice Finding
di: Xuan, Xiwei, et al.
Pubblicazione: (2024) -
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
di: Panigrahi, Indu, et al.
Pubblicazione: (2025) -
InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
di: Lin, Yukang, et al.
Pubblicazione: (2025)