Eagle: Ethical Dataset Given from Real Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Kaneko, Masahiro, Bollegala, Danushka, Baldwin, Timothy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
In-Contextual Gender Bias Suppression for Large Language Models
by: Oba, Daisuke, et al.
Published: (2023)
by: Oba, Daisuke, et al.
Published: (2023)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
A Semantic Distance Metric Learning approach for Lexical Semantic Change Detection
by: Aida, Taichi, et al.
Published: (2024)
by: Aida, Taichi, et al.
Published: (2024)
Investigating the Contextualised Word Embedding Dimensions Specified for Contextual and Temporal Semantic Changes
by: Aida, Taichi, et al.
Published: (2024)
by: Aida, Taichi, et al.
Published: (2024)
SCDTour: Embedding Axis Ordering and Merging for Interpretable Semantic Change Detection
by: Aida, Taichi, et al.
Published: (2025)
by: Aida, Taichi, et al.
Published: (2025)
Map of Encoders -- Mapping Sentence Encoders using Quantum Relative Entropy
by: Zhang, Gaifan, et al.
Published: (2026)
by: Zhang, Gaifan, et al.
Published: (2026)
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
by: Yamamoto, Taisei, et al.
Published: (2025)
by: Yamamoto, Taisei, et al.
Published: (2025)
Evaluating Unsupervised Dimensionality Reduction Methods for Pretrained Sentence Embeddings
by: Zhang, Gaifan, et al.
Published: (2024)
by: Zhang, Gaifan, et al.
Published: (2024)
Improving Diversity of Commonsense Generation by Large Language Models via In-Context Learning
by: Zhang, Tianhui, et al.
Published: (2024)
by: Zhang, Tianhui, et al.
Published: (2024)
Evaluating the Effect of Retrieval Augmentation on Social Biases
by: Zhang, Tianhui, et al.
Published: (2025)
by: Zhang, Tianhui, et al.
Published: (2025)
Synthetic Data Generation for Training Diversified Commonsense Reasoning Models
by: Zhang, Tianhui, et al.
Published: (2026)
by: Zhang, Tianhui, et al.
Published: (2026)
Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models
by: Zhang, Gaifan, et al.
Published: (2025)
by: Zhang, Gaifan, et al.
Published: (2025)
CASE -- Condition-Aware Sentence Embeddings for Conditional Semantic Textual Similarity Measurement
by: Zhang, Gaifan, et al.
Published: (2025)
by: Zhang, Gaifan, et al.
Published: (2025)
Evaluating the Evaluation of Diversity in Commonsense Generation
by: Zhang, Tianhui, et al.
Published: (2025)
by: Zhang, Tianhui, et al.
Published: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language Models
by: Zhou, Yi, et al.
Published: (2024)
by: Zhou, Yi, et al.
Published: (2024)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
by: Stuart, Harry, et al.
Published: (2026)
by: Stuart, Harry, et al.
Published: (2026)
Improving Unsupervised Constituency Parsing via Maximizing Semantic Information
by: Chen, Junjie, et al.
Published: (2024)
by: Chen, Junjie, et al.
Published: (2024)
Unsupervised Parsing by Searching for Frequent Word Sequences among Sentences with Equivalent Predicate-Argument Structures
by: Chen, Junjie, et al.
Published: (2024)
by: Chen, Junjie, et al.
Published: (2024)
Neuron-Level Analysis of Cultural Understanding in Large Language Models
by: Yamamoto, Taisei, et al.
Published: (2025)
by: Yamamoto, Taisei, et al.
Published: (2025)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
by: Kaneko, Masahiro, et al.
Published: (2026)
by: Kaneko, Masahiro, et al.
Published: (2026)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Improving Pre-trained Language Model Sensitivity via Mask Specific losses: A case study on Biomedical NER
by: Abaho, Micheal, et al.
Published: (2024)
by: Abaho, Micheal, et al.
Published: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
Arabic Dataset for LLM Safeguard Evaluation
by: Ashraf, Yasser, et al.
Published: (2024)
by: Ashraf, Yasser, et al.
Published: (2024)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
On the Alignment of Large Language Models with Global Human Opinion
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
by: Kaneko, Masahiro, et al.
Published: (2023)
by: Kaneko, Masahiro, et al.
Published: (2023)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
by: Loem, Mengsay, et al.
Published: (2023)
by: Loem, Mengsay, et al.
Published: (2023)
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners?
by: Kuribayashi, Tatsuki, et al.
Published: (2023)
by: Kuribayashi, Tatsuki, et al.
Published: (2023)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
Similar Items
-
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
by: Kaneko, Masahiro, et al.
Published: (2024) -
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
by: Kaneko, Masahiro, et al.
Published: (2024) -
In-Contextual Gender Bias Suppression for Large Language Models
by: Oba, Daisuke, et al.
Published: (2023) -
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026) -
A Semantic Distance Metric Learning approach for Lexical Semantic Change Detection
by: Aida, Taichi, et al.
Published: (2024)