Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Budagam, Devichand, Kumar, Ashutosh, Khoshnoodi, Mahsa, KJ, Sankalp, Jain, Vinija, Chadha, Aman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
von: Chauhan, Anay, et al.
Veröffentlicht: (2026)
von: Chauhan, Anay, et al.
Veröffentlicht: (2026)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
von: Peláez-González, Carlos, et al.
Veröffentlicht: (2025)
von: Peláez-González, Carlos, et al.
Veröffentlicht: (2025)
APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification
von: Chernodub, Artem, et al.
Veröffentlicht: (2025)
von: Chernodub, Artem, et al.
Veröffentlicht: (2025)
Building and Aligning Comparable Corpora
von: Saad, Motaz, et al.
Veröffentlicht: (2025)
von: Saad, Motaz, et al.
Veröffentlicht: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks
von: Figueiredo, Vanessa
Veröffentlicht: (2025)
von: Figueiredo, Vanessa
Veröffentlicht: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
Evaluating Steering Techniques using Human Similarity Judgments
von: Studdiford, Zach, et al.
Veröffentlicht: (2025)
von: Studdiford, Zach, et al.
Veröffentlicht: (2025)
Aligning the Norwegian UD Treebank with Entity and Coreference Information
von: Jørgensen, Tollef Emil, et al.
Veröffentlicht: (2023)
von: Jørgensen, Tollef Emil, et al.
Veröffentlicht: (2023)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
von: Pathak, Gangesh, et al.
Veröffentlicht: (2025)
von: Pathak, Gangesh, et al.
Veröffentlicht: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
von: The Omnilingual MT Team, et al.
Veröffentlicht: (2025)
von: The Omnilingual MT Team, et al.
Veröffentlicht: (2025)
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
Less Is More: Cognitive Load and the Single-Prompt Ceiling in LLM Mathematical Reasoning
von: Cazares, Manuel Israel
Veröffentlicht: (2026)
von: Cazares, Manuel Israel
Veröffentlicht: (2026)
HiPS: Hierarchical PDF Segmentation of Textbooks
von: Wehnert, Sabine, et al.
Veröffentlicht: (2025)
von: Wehnert, Sabine, et al.
Veröffentlicht: (2025)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
von: Lübbers, Christopher Lee
Veröffentlicht: (2025)
von: Lübbers, Christopher Lee
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
von: Bandooni, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bandooni, Ashutosh, et al.
Veröffentlicht: (2025)
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening
von: Lorenzoni, Giuliano, et al.
Veröffentlicht: (2026)
von: Lorenzoni, Giuliano, et al.
Veröffentlicht: (2026)
Evaluating 5W3H Structured Prompting for Intent Alignment in Human-AI Interaction
von: Gang, Peng
Veröffentlicht: (2026)
von: Gang, Peng
Veröffentlicht: (2026)
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)
IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages
von: VP, Sumesh
Veröffentlicht: (2026)
von: VP, Sumesh
Veröffentlicht: (2026)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
von: Forchheimer, Robert
Veröffentlicht: (2026)
von: Forchheimer, Robert
Veröffentlicht: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
von: Nieth, Björn, et al.
Veröffentlicht: (2026)
von: Nieth, Björn, et al.
Veröffentlicht: (2026)
Active Few-Shot Learning for Text Classification
von: Ahmadnia, Saeed, et al.
Veröffentlicht: (2025)
von: Ahmadnia, Saeed, et al.
Veröffentlicht: (2025)
ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
von: Yi, Seungyoun, et al.
Veröffentlicht: (2025)
von: Yi, Seungyoun, et al.
Veröffentlicht: (2025)
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
von: Stewart, Ian, et al.
Veröffentlicht: (2024)
von: Stewart, Ian, et al.
Veröffentlicht: (2024)
PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI
von: Chaitanya, Keshava, et al.
Veröffentlicht: (2026)
von: Chaitanya, Keshava, et al.
Veröffentlicht: (2026)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
von: Velampalli, Sirisha, et al.
Veröffentlicht: (2025)
von: Velampalli, Sirisha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
von: Chauhan, Anay, et al.
Veröffentlicht: (2026) -
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
von: Khanna, Danush, et al.
Veröffentlicht: (2025) -
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025) -
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
von: Wang, Xintao, et al.
Veröffentlicht: (2026) -
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)