Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yongkang, Feng, Shi, Wang, Daling, Zhang, Yifei, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChatZero:Zero-shot Cross-Lingual Dialogue Generation via Pseudo-Target Language
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
A Unified Data Augmentation Framework for Low-Resource Multi-Domain Dialogue Generation
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
by: Wang, Zijing, et al.
Published: (2026)
by: Wang, Zijing, et al.
Published: (2026)
DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging
by: Wang, Zijing, et al.
Published: (2026)
by: Wang, Zijing, et al.
Published: (2026)
High-Rank Structured Modulation for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction
by: Mu, Shiyi, et al.
Published: (2025)
by: Mu, Shiyi, et al.
Published: (2025)
SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
by: Kennedy, Molly, et al.
Published: (2026)
by: Kennedy, Molly, et al.
Published: (2026)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
TOOL-ED: Enhancing Empathetic Response Generation with the Tool Calling Capability of LLM
by: Cao, Huiying, et al.
Published: (2024)
by: Cao, Huiying, et al.
Published: (2024)
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
NEAT: Neuron-Based Early Exit for Large Reasoning Models
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
RoCar: A Relationship Network-based Evaluation Method for Large Language Models
by: Wang, Ming, et al.
Published: (2023)
by: Wang, Ming, et al.
Published: (2023)
You Can't Fight in Here! This is BBS!
by: Futrell, Richard, et al.
Published: (2026)
by: Futrell, Richard, et al.
Published: (2026)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
by: Cai, Dexian, et al.
Published: (2025)
by: Cai, Dexian, et al.
Published: (2025)
MoLAN: A Unified Modality-Aware Noise Dynamic Editing Framework for Multimodal Sentiment Analysis
by: Xu, Xingle, et al.
Published: (2025)
by: Xu, Xingle, et al.
Published: (2025)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
STICKERCONV: Generating Multimodal Empathetic Responses from Scratch
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
by: Veloso, Leonor, et al.
Published: (2026)
by: Veloso, Leonor, et al.
Published: (2026)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
Generative Emotion Cause Explanation in Multimodal Conversations
by: Wang, Lin, et al.
Published: (2024)
by: Wang, Lin, et al.
Published: (2024)
Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering
by: Zhang, Xiaoming, et al.
Published: (2024)
by: Zhang, Xiaoming, et al.
Published: (2024)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
by: Veitsman, Yana, et al.
Published: (2026)
by: Veitsman, Yana, et al.
Published: (2026)
Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment
by: Xhelili, Orgest, et al.
Published: (2024)
by: Xhelili, Orgest, et al.
Published: (2024)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
by: Gerstner, Sebastian, et al.
Published: (2026)
by: Gerstner, Sebastian, et al.
Published: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
by: Gerstner, Sebastian, et al.
Published: (2025)
by: Gerstner, Sebastian, et al.
Published: (2025)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
by: Mortensen, David R., et al.
Published: (2024)
by: Mortensen, David R., et al.
Published: (2024)
AbsenceBench: Language Models Can't Tell What's Missing
by: Fu, Harvey Yiyun, et al.
Published: (2025)
by: Fu, Harvey Yiyun, et al.
Published: (2025)
GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification
by: Wan, Faxian, et al.
Published: (2026)
by: Wan, Faxian, et al.
Published: (2026)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
by: Liu, Yihong, et al.
Published: (2023)
by: Liu, Yihong, et al.
Published: (2023)
On the Entity-Level Alignment in Crosslingual Consistency
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
by: Feng, Qi, et al.
Published: (2025)
by: Feng, Qi, et al.
Published: (2025)
Relational Linearity is a Predictor of Hallucinations
by: Lu, Yuetian, et al.
Published: (2026)
by: Lu, Yuetian, et al.
Published: (2026)
Similar Items
-
ChatZero:Zero-shot Cross-Lingual Dialogue Generation via Pseudo-Target Language
by: Liu, Yongkang, et al.
Published: (2024) -
A Unified Data Augmentation Framework for Low-Resource Multi-Domain Dialogue Generation
by: Liu, Yongkang, et al.
Published: (2024) -
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy
by: Liu, Yongkang, et al.
Published: (2024) -
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset
by: Liu, Yongkang, et al.
Published: (2026) -
PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
by: Wang, Zijing, et al.
Published: (2026)