ESG Classification by Implicit Rule Learning via GPT-4
Fuente:
arXiv
Salvato in:
| Autori principali: | Yun, Hyo Jeong, Kim, Chanyoung, Hahm, Moonjeong, Kim, Kyuri, Son, Guijin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Multi-Step Reasoning in Korean and the Emergent Mirage
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
di: Choi, Dasol, et al.
Pubblicazione: (2024)
di: Choi, Dasol, et al.
Pubblicazione: (2024)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
Won: Establishing Best Practices for Korean Financial NLP
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
KAIO: A Collection of More Challenging Korean Questions
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Controlling Language Confusion in Multilingual LLMs
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
di: Son, Guijin, et al.
Pubblicazione: (2023)
di: Son, Guijin, et al.
Pubblicazione: (2023)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
di: Yun, Jungmin, et al.
Pubblicazione: (2024)
di: Yun, Jungmin, et al.
Pubblicazione: (2024)
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
di: Im, Kyuri, et al.
Pubblicazione: (2026)
di: Im, Kyuri, et al.
Pubblicazione: (2026)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
di: Lee, Hanwool, et al.
Pubblicazione: (2025)
di: Lee, Hanwool, et al.
Pubblicazione: (2025)
Automated Information Extraction from Thyroid Operation Narrative: A Comparative Study of GPT-4 and Fine-tuned KoELECTRA
di: Jang, Dongsuk, et al.
Pubblicazione: (2024)
di: Jang, Dongsuk, et al.
Pubblicazione: (2024)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
di: Hahm, Sungeun, et al.
Pubblicazione: (2025)
di: Hahm, Sungeun, et al.
Pubblicazione: (2025)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
di: Kim, Yoonsu, et al.
Pubblicazione: (2024)
di: Kim, Yoonsu, et al.
Pubblicazione: (2024)
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
di: Jeon, Minkyeong, et al.
Pubblicazione: (2025)
di: Jeon, Minkyeong, et al.
Pubblicazione: (2025)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
di: Liu, Cheng, et al.
Pubblicazione: (2025)
di: Liu, Cheng, et al.
Pubblicazione: (2025)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
Bayesian Multi-Task Transfer Learning for Soft Prompt Tuning
di: Lee, Haeju, et al.
Pubblicazione: (2024)
di: Lee, Haeju, et al.
Pubblicazione: (2024)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
di: Kim, Wonjoong, et al.
Pubblicazione: (2025)
di: Kim, Wonjoong, et al.
Pubblicazione: (2025)
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
di: Lee, Seungbeen, et al.
Pubblicazione: (2025)
di: Lee, Seungbeen, et al.
Pubblicazione: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules
di: Chen, Kejie, et al.
Pubblicazione: (2024)
di: Chen, Kejie, et al.
Pubblicazione: (2024)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering
di: Kim, Byeolhee, et al.
Pubblicazione: (2026)
di: Kim, Byeolhee, et al.
Pubblicazione: (2026)
On the Robustness of Reward Models for Language Model Alignment
di: Hong, Jiwoo, et al.
Pubblicazione: (2025)
di: Hong, Jiwoo, et al.
Pubblicazione: (2025)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
di: Kim, Seungju, et al.
Pubblicazione: (2024)
di: Kim, Seungju, et al.
Pubblicazione: (2024)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
di: Choi, Dasol, et al.
Pubblicazione: (2025) -
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
di: Son, Guijin, et al.
Pubblicazione: (2024) -
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
di: Gwak, Minju, et al.
Pubblicazione: (2025) -
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025) -
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
di: Lee, Nahyun, et al.
Pubblicazione: (2026)