LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yi, Duan, Hanyu, Liu, Jiaxin, Tam, Kar Yan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
by: Duan, Hanyu, et al.
Published: (2024)
by: Duan, Hanyu, et al.
Published: (2024)
Beyond Surface Similarity: Detecting Subtle Semantic Shifts in Financial Narratives
by: Liu, Jiaxin, et al.
Published: (2024)
by: Liu, Jiaxin, et al.
Published: (2024)
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads
by: Yang, Yi, et al.
Published: (2023)
by: Yang, Yi, et al.
Published: (2023)
Evaluating and Aligning Human Economic Risk Preferences in LLMs
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
by: Jiang, Jingzhou, et al.
Published: (2026)
by: Jiang, Jingzhou, et al.
Published: (2026)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
by: Deng, Ningyuan, et al.
Published: (2025)
by: Deng, Ningyuan, et al.
Published: (2025)
FLARE: Task-agnostic embedding model evaluation through a normalization process
by: Jiang, Jingzhou, et al.
Published: (2026)
by: Jiang, Jingzhou, et al.
Published: (2026)
Measuring Scalar Constructs in Social Science with LLMs
by: Licht, Hauke, et al.
Published: (2025)
by: Licht, Hauke, et al.
Published: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
by: Gupta, Ashim, et al.
Published: (2025)
by: Gupta, Ashim, et al.
Published: (2025)
REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
by: Hu, Chuxuan, et al.
Published: (2025)
by: Hu, Chuxuan, et al.
Published: (2025)
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
by: Yang, Lingyi, et al.
Published: (2023)
by: Yang, Lingyi, et al.
Published: (2023)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
by: Lorandi, Michela, et al.
Published: (2024)
by: Lorandi, Michela, et al.
Published: (2024)
Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches
by: Shah, Syed Mehtab Hussain, et al.
Published: (2026)
by: Shah, Syed Mehtab Hussain, et al.
Published: (2026)
Robust Evaluation Measures for Evaluating Social Biases in Masked Language Models
by: Liu, Yang
Published: (2024)
by: Liu, Yang
Published: (2024)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
by: Lv, Kai, et al.
Published: (2024)
by: Lv, Kai, et al.
Published: (2024)
SQLFixAgent: Towards Semantic-Accurate Text-to-SQL Parsing via Consistency-Enhanced Multi-Agent Collaboration
by: Cen, Jipeng, et al.
Published: (2024)
by: Cen, Jipeng, et al.
Published: (2024)
Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs
by: Kim, Taejin, et al.
Published: (2025)
by: Kim, Taejin, et al.
Published: (2025)
Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
by: Cheng, Zihao, et al.
Published: (2024)
by: Cheng, Zihao, et al.
Published: (2024)
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Measuring AI "Slop" in Text
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
by: Li, Baishi, et al.
Published: (2026)
by: Li, Baishi, et al.
Published: (2026)
Generic Embedding-Based Lexicons for Transparent and Reproducible Text Scoring
by: Moez, Catherine
Published: (2024)
by: Moez, Catherine
Published: (2024)
Research on Graph-Retrieval Augmented Generation Based on Historical Text Knowledge Graphs
by: Fan, Yang, et al.
Published: (2025)
by: Fan, Yang, et al.
Published: (2025)
Measuring Psychological Depth in Language Models
by: Harel-Canada, Fabrice, et al.
Published: (2024)
by: Harel-Canada, Fabrice, et al.
Published: (2024)
Measurement of LLM's Philosophies of Human Nature
by: Ni, Minheng, et al.
Published: (2025)
by: Ni, Minheng, et al.
Published: (2025)
Is This Collection Worth My LLM's Time? Automatically Measuring Information Potential in Text Corpora
by: Karch, Tristan, et al.
Published: (2025)
by: Karch, Tristan, et al.
Published: (2025)
Arithmetic Reasoning with LLM: Prolog Generation & Permutation
by: Yang, Xiaocheng, et al.
Published: (2024)
by: Yang, Xiaocheng, et al.
Published: (2024)
SearchLLM: Detecting LLM Paraphrased Text by Measuring the Similarity with Regeneration of the Candidate Source via Search Engine
by: Nguyen-Son, Hoang-Quoc, et al.
Published: (2026)
by: Nguyen-Son, Hoang-Quoc, et al.
Published: (2026)
Measuring Mental Health Variables in Computational Research: Toward Validated, Dimensional, and Transdiagnostic Approaches
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
CPTuning: Contrastive Prompt Tuning for Generative Relation Extraction
by: Duan, Jiaxin, et al.
Published: (2025)
by: Duan, Jiaxin, et al.
Published: (2025)
Trustworthy Social Bias Measurement
by: Bommasani, Rishi, et al.
Published: (2022)
by: Bommasani, Rishi, et al.
Published: (2022)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning
by: Wang, Rongsheng, et al.
Published: (2024)
by: Wang, Rongsheng, et al.
Published: (2024)
Navigating the Prompt Space: Improving LLM Classification of Social Science Texts Through Prompt Engineering
by: Gunes, Erkan, et al.
Published: (2026)
by: Gunes, Erkan, et al.
Published: (2026)
Robust Predictive Modeling Under Unseen Data Distribution Shifts: A Methodological Commentary
by: Duan, Hanyu, et al.
Published: (2025)
by: Duan, Hanyu, et al.
Published: (2025)
Ready2Unlearn: A Learning-Time Approach for Preparing Models with Future Unlearning Readiness
by: Duan, Hanyu, et al.
Published: (2025)
by: Duan, Hanyu, et al.
Published: (2025)
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
by: Gao, Bufan, et al.
Published: (2025)
by: Gao, Bufan, et al.
Published: (2025)
Similar Items
-
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
by: Duan, Hanyu, et al.
Published: (2024) -
Beyond Surface Similarity: Detecting Subtle Semantic Shifts in Financial Narratives
by: Liu, Jiaxin, et al.
Published: (2024) -
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads
by: Yang, Yi, et al.
Published: (2023) -
Evaluating and Aligning Human Economic Risk Preferences in LLMs
by: Liu, Jiaxin, et al.
Published: (2025) -
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
by: Jiang, Jingzhou, et al.
Published: (2026)