ESG Classification by Implicit Rule Learning via GPT-4
Fuente:
arXiv
Guardado en:
| Autores principales: | Yun, Hyo Jeong, Kim, Chanyoung, Hahm, Moonjeong, Kim, Kyuri, Son, Guijin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
por: Choi, Dasol, et al.
Publicado: (2025)
por: Choi, Dasol, et al.
Publicado: (2025)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
por: Gwak, Minju, et al.
Publicado: (2025)
por: Gwak, Minju, et al.
Publicado: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
por: Gwak, Minju, et al.
Publicado: (2025)
por: Gwak, Minju, et al.
Publicado: (2025)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
por: Lee, Nahyun, et al.
Publicado: (2026)
por: Lee, Nahyun, et al.
Publicado: (2026)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
por: Lee, Nahyun, et al.
Publicado: (2026)
por: Lee, Nahyun, et al.
Publicado: (2026)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Multi-Step Reasoning in Korean and the Emergent Mirage
por: Son, Guijin, et al.
Publicado: (2025)
por: Son, Guijin, et al.
Publicado: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
por: Ko, Hyunwoo, et al.
Publicado: (2025)
por: Ko, Hyunwoo, et al.
Publicado: (2025)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
por: Choi, Dasol, et al.
Publicado: (2024)
por: Choi, Dasol, et al.
Publicado: (2024)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
por: Hong, Seokhee, et al.
Publicado: (2025)
por: Hong, Seokhee, et al.
Publicado: (2025)
Won: Establishing Best Practices for Korean Financial NLP
por: Son, Guijin, et al.
Publicado: (2025)
por: Son, Guijin, et al.
Publicado: (2025)
KAIO: A Collection of More Challenging Korean Questions
por: Lee, Nahyun, et al.
Publicado: (2025)
por: Lee, Nahyun, et al.
Publicado: (2025)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
por: Son, Guijin, et al.
Publicado: (2025)
por: Son, Guijin, et al.
Publicado: (2025)
Controlling Language Confusion in Multilingual LLMs
por: Lee, Nahyun, et al.
Publicado: (2025)
por: Lee, Nahyun, et al.
Publicado: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
por: Son, Guijin, et al.
Publicado: (2023)
por: Son, Guijin, et al.
Publicado: (2023)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
por: Yun, Jungmin, et al.
Publicado: (2024)
por: Yun, Jungmin, et al.
Publicado: (2024)
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
por: Son, Guijin, et al.
Publicado: (2026)
por: Son, Guijin, et al.
Publicado: (2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
por: Kim, Eunsu, et al.
Publicado: (2025)
por: Kim, Eunsu, et al.
Publicado: (2025)
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
por: Im, Kyuri, et al.
Publicado: (2026)
por: Im, Kyuri, et al.
Publicado: (2026)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
por: Lee, Hanwool, et al.
Publicado: (2025)
por: Lee, Hanwool, et al.
Publicado: (2025)
Automated Information Extraction from Thyroid Operation Narrative: A Comparative Study of GPT-4 and Fine-tuned KoELECTRA
por: Jang, Dongsuk, et al.
Publicado: (2024)
por: Jang, Dongsuk, et al.
Publicado: (2024)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
por: Hahm, Sungeun, et al.
Publicado: (2025)
por: Hahm, Sungeun, et al.
Publicado: (2025)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
por: In, Yeonjun, et al.
Publicado: (2025)
por: In, Yeonjun, et al.
Publicado: (2025)
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
por: Son, Guijin, et al.
Publicado: (2026)
por: Son, Guijin, et al.
Publicado: (2026)
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
por: Kim, Yoonsu, et al.
Publicado: (2024)
por: Kim, Yoonsu, et al.
Publicado: (2024)
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
por: Jeon, Minkyeong, et al.
Publicado: (2025)
por: Jeon, Minkyeong, et al.
Publicado: (2025)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
por: Liu, Cheng, et al.
Publicado: (2025)
por: Liu, Cheng, et al.
Publicado: (2025)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
por: In, Yeonjun, et al.
Publicado: (2025)
por: In, Yeonjun, et al.
Publicado: (2025)
Bayesian Multi-Task Transfer Learning for Soft Prompt Tuning
por: Lee, Haeju, et al.
Publicado: (2024)
por: Lee, Haeju, et al.
Publicado: (2024)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
por: Kim, Wonjoong, et al.
Publicado: (2025)
por: Kim, Wonjoong, et al.
Publicado: (2025)
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
por: Lee, Seungbeen, et al.
Publicado: (2025)
por: Lee, Seungbeen, et al.
Publicado: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
por: Lee, Nahyun, et al.
Publicado: (2026)
por: Lee, Nahyun, et al.
Publicado: (2026)
MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules
por: Chen, Kejie, et al.
Publicado: (2024)
por: Chen, Kejie, et al.
Publicado: (2024)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering
por: Kim, Byeolhee, et al.
Publicado: (2026)
por: Kim, Byeolhee, et al.
Publicado: (2026)
On the Robustness of Reward Models for Language Model Alignment
por: Hong, Jiwoo, et al.
Publicado: (2025)
por: Hong, Jiwoo, et al.
Publicado: (2025)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
por: Kim, Seungju, et al.
Publicado: (2024)
por: Kim, Seungju, et al.
Publicado: (2024)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Ejemplares similares
-
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
por: Choi, Dasol, et al.
Publicado: (2025) -
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
por: Son, Guijin, et al.
Publicado: (2024) -
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
por: Gwak, Minju, et al.
Publicado: (2025) -
Revisiting the UID Hypothesis in LLM Reasoning Traces
por: Gwak, Minju, et al.
Publicado: (2025) -
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
por: Lee, Nahyun, et al.
Publicado: (2026)