LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kahng, Minsuk, Tenney, Ian, Pushkarna, Mahima, Liu, Michael Xieyang, Wexler, James, Reif, Emily, Kallarackal, Krystal, Chang, Minsuk, Terry, Michael, Dixon, Lucas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automatic Histograms: Leveraging Language Models for Text Dataset Exploration
von: Reif, Emily, et al.
Veröffentlicht: (2024)
von: Reif, Emily, et al.
Veröffentlicht: (2024)
Understanding the Dataset Practitioners Behind Large Language Model Development
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
The Evolution of LLM Adoption in Industry Data Curation Practices
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
Interactive Prompt Debugging with Sequence Salience
von: Tenney, Ian, et al.
Veröffentlicht: (2024)
von: Tenney, Ian, et al.
Veröffentlicht: (2024)
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
von: Jung, Minji, et al.
Veröffentlicht: (2026)
von: Jung, Minji, et al.
Veröffentlicht: (2026)
Designing for Human-Agent Alignment: Understanding what humans want from their agents
von: Goyal, Nitesh, et al.
Veröffentlicht: (2024)
von: Goyal, Nitesh, et al.
Veröffentlicht: (2024)
VLSlice: Interactive Vision-and-Language Slice Discovery
von: Slyman, Eric, et al.
Veröffentlicht: (2023)
von: Slyman, Eric, et al.
Veröffentlicht: (2023)
From Perception to Decision: Assessing the Role of Chart Types Affordances in High-Level Decision Tasks
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
"We Need Structured Output": Towards User-centered Constraints on Large Language Model Output
von: Liu, Michael Xieyang, et al.
Veröffentlicht: (2024)
von: Liu, Michael Xieyang, et al.
Veröffentlicht: (2024)
Grid Labeling: Crowdsourcing Task-Specific Importance from Visualizations
von: Chang, Minsuk, et al.
Veröffentlicht: (2025)
von: Chang, Minsuk, et al.
Veröffentlicht: (2025)
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
von: Chang, Minsuk, et al.
Veröffentlicht: (2025)
von: Chang, Minsuk, et al.
Veröffentlicht: (2025)
Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness
von: Lee, Soohyun, et al.
Veröffentlicht: (2024)
von: Lee, Soohyun, et al.
Veröffentlicht: (2024)
Efficiently Crowdsourcing Visual Importance with Punch-Hole Annotation
von: Chang, Minsuk, et al.
Veröffentlicht: (2024)
von: Chang, Minsuk, et al.
Veröffentlicht: (2024)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
von: Ibrahim, Lujain, et al.
Veröffentlicht: (2025)
von: Ibrahim, Lujain, et al.
Veröffentlicht: (2025)
Compass vs Railway Tracks: Unpacking User Mental Models for Communicating Long-Horizon Work to Humans vs. AI
von: Petridis, Savvas, et al.
Veröffentlicht: (2026)
von: Petridis, Savvas, et al.
Veröffentlicht: (2026)
Extracting Human Attention through Crowdsourced Patch Labeling
von: Chang, Minsuk, et al.
Veröffentlicht: (2024)
von: Chang, Minsuk, et al.
Veröffentlicht: (2024)
Interactive AI Alignment: Specification, Process, and Evaluation Alignment
von: Terry, Michael, et al.
Veröffentlicht: (2023)
von: Terry, Michael, et al.
Veröffentlicht: (2023)
CHOMP: Multimodal Chewing Side Detection with Earphones
von: Hummel, Jonas, et al.
Veröffentlicht: (2026)
von: Hummel, Jonas, et al.
Veröffentlicht: (2026)
Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning
von: Liu, Michael Xieyang, et al.
Veröffentlicht: (2025)
von: Liu, Michael Xieyang, et al.
Veröffentlicht: (2025)
Believing Anthropomorphism: Examining the Role of Anthropomorphic Cues on Trust in Large Language Models
von: Cohn, Michelle, et al.
Veröffentlicht: (2024)
von: Cohn, Michelle, et al.
Veröffentlicht: (2024)
Tasks, Time, and Tools: Quantifying Online Sensemaking Efforts Through a Survey-based Study
von: Kuznetsov, Andrew, et al.
Veröffentlicht: (2024)
von: Kuznetsov, Andrew, et al.
Veröffentlicht: (2024)
In Situ AI Prototyping: Infusing Multimodal Prompts into Mobile Settings with MobileMaker
von: Petridis, Savvas, et al.
Veröffentlicht: (2024)
von: Petridis, Savvas, et al.
Veröffentlicht: (2024)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies
von: Maldaner, Matheus Kunzler, et al.
Veröffentlicht: (2025)
von: Maldaner, Matheus Kunzler, et al.
Veröffentlicht: (2025)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
A Contextual Inquiry of People with Vision Impairments in Cooking
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2024)
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2024)
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
von: Qian, Crystal, et al.
Veröffentlicht: (2024)
Exploring the Capability of LLMs in Performing Low-Level Visual Analytic Tasks on SVG Data Visualizations
von: Xu, Zhongzheng, et al.
Veröffentlicht: (2024)
von: Xu, Zhongzheng, et al.
Veröffentlicht: (2024)
ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows
von: Andrienko, Natalia, et al.
Veröffentlicht: (2026)
von: Andrienko, Natalia, et al.
Veröffentlicht: (2026)
Show Me Your Best Side: Characteristics of User-Preferred Perspectives for 3D Graph Drawings
von: Joos, Lucas, et al.
Veröffentlicht: (2025)
von: Joos, Lucas, et al.
Veröffentlicht: (2025)
RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
von: Tian, Haoyu, et al.
Veröffentlicht: (2026)
von: Tian, Haoyu, et al.
Veröffentlicht: (2026)
BrainDecoder: Style-Based Visual Decoding of EEG Signals
von: Choi, Minsuk, et al.
Veröffentlicht: (2024)
von: Choi, Minsuk, et al.
Veröffentlicht: (2024)
Ambient Analytics: Calm Technology for Immersive Visualization and Sensemaking
von: Hubenschmid, Sebastian, et al.
Veröffentlicht: (2026)
von: Hubenschmid, Sebastian, et al.
Veröffentlicht: (2026)
AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages
von: Wu, Muzhe, et al.
Veröffentlicht: (2025)
von: Wu, Muzhe, et al.
Veröffentlicht: (2025)
A Comparative Visual Analytics Framework for Evaluating Evolutionary Processes in Multi-objective Optimization
von: Huang, Yansong, et al.
Veröffentlicht: (2023)
von: Huang, Yansong, et al.
Veröffentlicht: (2023)
Comparative evaluation of the web-based contiguous cartogram generation tool go-cart.io
von: Duncan, Ian K., et al.
Veröffentlicht: (2022)
von: Duncan, Ian K., et al.
Veröffentlicht: (2022)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
von: Rastogi, Charvi, et al.
Veröffentlicht: (2026)
von: Rastogi, Charvi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Automatic Histograms: Leveraging Language Models for Text Dataset Exploration
von: Reif, Emily, et al.
Veröffentlicht: (2024) -
Understanding the Dataset Practitioners Behind Large Language Model Development
von: Qian, Crystal, et al.
Veröffentlicht: (2024) -
The Evolution of LLM Adoption in Industry Data Curation Practices
von: Qian, Crystal, et al.
Veröffentlicht: (2024) -
Interactive Prompt Debugging with Sequence Salience
von: Tenney, Ian, et al.
Veröffentlicht: (2024) -
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
von: Lee, Minjae, et al.
Veröffentlicht: (2025)