Saved in:
| Main Authors: | Brunello, Nicolò, Rigamonti, Davide, Sassella, Andrea, Scotti, Vincenzo, Carman, Mark James |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.13858 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026)
by: Toschi, Federico, et al.
Published: (2026)
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
by: Singh, Raul, et al.
Published: (2025)
by: Singh, Raul, et al.
Published: (2025)
Are complicated loss functions necessary for teaching LLMs to reason?
by: Carrino, Gabriele, et al.
Published: (2026)
by: Carrino, Gabriele, et al.
Published: (2026)
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
by: Sassella, Andrea, et al.
Published: (2026)
by: Sassella, Andrea, et al.
Published: (2026)
Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
by: Nguyen, Hy, et al.
Published: (2024)
by: Nguyen, Hy, et al.
Published: (2024)
More Church-Rosser Proofs in BELUGA
by: Momigliano, Alberto, et al.
Published: (2024)
by: Momigliano, Alberto, et al.
Published: (2024)
A Shared Geometry of Difficulty in Multilingual Language Models
by: Civelli, Stefano, et al.
Published: (2026)
by: Civelli, Stefano, et al.
Published: (2026)
Expect the unexpected: Harnessing Sentence Completion for Sarcasm Detection
by: Joshi, Aditya, et al.
Published: (2017)
by: Joshi, Aditya, et al.
Published: (2017)
Automata-less Monitoring via Trace-Checking (Extended Version)
by: Brunello, Andrea, et al.
Published: (2025)
by: Brunello, Andrea, et al.
Published: (2025)
Do LLMs Really Struggle at NL-FOL Translation? Revealing their Strengths via a Novel Benchmarking Strategy
by: Brunello, Andrea, et al.
Published: (2025)
by: Brunello, Andrea, et al.
Published: (2025)
ReasonGraph: Visualisation of Reasoning Paths
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
Do You Understand How I Feel?: Towards Verified Empathy in Therapy Chatbots
by: Dettori, Francesco, et al.
Published: (2026)
by: Dettori, Francesco, et al.
Published: (2026)
Interpretable Early Failure Detection via Machine Learning and Trace Checking-based Monitoring
by: Brunello, Andrea, et al.
Published: (2025)
by: Brunello, Andrea, et al.
Published: (2025)
SumTra: A Differentiable Pipeline for Few-Shot Cross-Lingual Summarization
by: Parnell, Jacob, et al.
Published: (2024)
by: Parnell, Jacob, et al.
Published: (2024)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
by: Corbo, Simone, et al.
Published: (2025)
by: Corbo, Simone, et al.
Published: (2025)
VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integration
by: Khaki, Saeed, et al.
Published: (2026)
by: Khaki, Saeed, et al.
Published: (2026)
Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models
by: Galatolo, Alessio, et al.
Published: (2025)
by: Galatolo, Alessio, et al.
Published: (2025)
LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models
by: Tufanov, Igor, et al.
Published: (2024)
by: Tufanov, Igor, et al.
Published: (2024)
Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging
by: Fabian, Thomas
Published: (2026)
by: Fabian, Thomas
Published: (2026)
DeVisE: Behavioral Testing of Medical Large Language Models
by: Tagliabue, Camila Zurdo, et al.
Published: (2025)
by: Tagliabue, Camila Zurdo, et al.
Published: (2025)
VisTW: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Visualising CTL Witnesses and Counterexamples -- Extended Version
by: Rensink, Arend
Published: (2026)
by: Rensink, Arend
Published: (2026)
VisCoder2: Building Multi-Language Visualization Coding Agents
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting
by: Gan, Shuyu, et al.
Published: (2025)
by: Gan, Shuyu, et al.
Published: (2025)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
by: Xie, Yupeng, et al.
Published: (2025)
by: Xie, Yupeng, et al.
Published: (2025)
On the Universal Truthfulness Hyperplane Inside LLMs
by: Liu, Junteng, et al.
Published: (2024)
by: Liu, Junteng, et al.
Published: (2024)
A Directed Graph Model and Experimental Framework for Design and Study of Time-Dependent Text Visualisation
by: Fan, Songhai, et al.
Published: (2026)
by: Fan, Songhai, et al.
Published: (2026)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
Text-to-TrajVis: Enabling Trajectory Data Visualizations from Natural Language Questions
by: Bai, Tian, et al.
Published: (2025)
by: Bai, Tian, et al.
Published: (2025)
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
by: Yue, Yuanhao, et al.
Published: (2026)
by: Yue, Yuanhao, et al.
Published: (2026)
Toolshed: Scale Tool-Equipped Agents with Advanced RAG-Tool Fusion and Tool Knowledge Bases
by: Lumer, Elias, et al.
Published: (2024)
by: Lumer, Elias, et al.
Published: (2024)
ChatVis: Automating Scientific Visualization with a Large Language Model
by: Mallick, Tanwi, et al.
Published: (2024)
by: Mallick, Tanwi, et al.
Published: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
by: Ascione, Grazia Sveva, et al.
Published: (2025)
by: Ascione, Grazia Sveva, et al.
Published: (2025)
Annotation Tool and Dataset for Fact-Checking Podcasts
by: Setty, Vinay, et al.
Published: (2025)
by: Setty, Vinay, et al.
Published: (2025)
Similar Items
-
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026) -
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
by: Singh, Raul, et al.
Published: (2025) -
Are complicated loss functions necessary for teaching LLMs to reason?
by: Carrino, Gabriele, et al.
Published: (2026) -
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
by: Sassella, Andrea, et al.
Published: (2026) -
Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
by: Nguyen, Hy, et al.
Published: (2024)