NJUST-KMG at TRAC-2024 Tasks 1 and 2: Offline Harm Potential Identification
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Jingyuan, Xu, Shengdong, Yang, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LegalLens Shared Task 2024: Legal Violation Identification in Unstructured Text
por: Hagag, Ben, et al.
Publicado: (2024)
por: Hagag, Ben, et al.
Publicado: (2024)
Investigating and Alleviating Harm Amplification in LLM Interactions
por: Guo, Ruohao, et al.
Publicado: (2026)
por: Guo, Ruohao, et al.
Publicado: (2026)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
por: Qi, Jingyuan, et al.
Publicado: (2023)
por: Qi, Jingyuan, et al.
Publicado: (2023)
Data Contamination Report from the 2024 CONDA Shared Task
por: Sainz, Oscar, et al.
Publicado: (2024)
por: Sainz, Oscar, et al.
Publicado: (2024)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
por: Chehbouni, Khaoula, et al.
Publicado: (2024)
por: Chehbouni, Khaoula, et al.
Publicado: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
por: Yan, Hao, et al.
Publicado: (2026)
por: Yan, Hao, et al.
Publicado: (2026)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
por: Hong, Yuzhong, et al.
Publicado: (2024)
por: Hong, Yuzhong, et al.
Publicado: (2024)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
por: Wang, Huaijie, et al.
Publicado: (2024)
por: Wang, Huaijie, et al.
Publicado: (2024)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
por: Stepanov, Ihor, et al.
Publicado: (2026)
por: Stepanov, Ihor, et al.
Publicado: (2026)
The Solution for The PST-KDD-2024 OAG-Challenge
por: Zhong, Shupeng, et al.
Publicado: (2024)
por: Zhong, Shupeng, et al.
Publicado: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
por: Wang, Shun, et al.
Publicado: (2025)
por: Wang, Shun, et al.
Publicado: (2025)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
por: Feng, Weitao, et al.
Publicado: (2025)
por: Feng, Weitao, et al.
Publicado: (2025)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
por: Lo, Chung-Hsiang, et al.
Publicado: (2026)
por: Lo, Chung-Hsiang, et al.
Publicado: (2026)
Representation Noising: A Defence Mechanism Against Harmful Finetuning
por: Rosati, Domenic, et al.
Publicado: (2024)
por: Rosati, Domenic, et al.
Publicado: (2024)
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
por: Yang, Ziyu, et al.
Publicado: (2026)
por: Yang, Ziyu, et al.
Publicado: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
por: Wu, Junda, et al.
Publicado: (2024)
por: Wu, Junda, et al.
Publicado: (2024)
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents
por: Zhu, Jialiang, et al.
Publicado: (2026)
por: Zhu, Jialiang, et al.
Publicado: (2026)
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
por: Wang, Shaobo, et al.
Publicado: (2026)
por: Wang, Shaobo, et al.
Publicado: (2026)
NeuroLoRA: Context-Aware Neuromodulation for Parameter-Efficient Multi-Task Adaptation
por: Yang, Yuxin, et al.
Publicado: (2026)
por: Yang, Yuxin, et al.
Publicado: (2026)
A Baseline for Self-state Identification and Classification in Mental Health Data: CLPsych 2025 Task
por: Kim, Laerdon
Publicado: (2025)
por: Kim, Laerdon
Publicado: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
por: Huang, Quzhe, et al.
Publicado: (2024)
por: Huang, Quzhe, et al.
Publicado: (2024)
Learning Task Representations from In-Context Learning
por: Saglam, Baturay, et al.
Publicado: (2025)
por: Saglam, Baturay, et al.
Publicado: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
por: Lee, Seanie, et al.
Publicado: (2024)
por: Lee, Seanie, et al.
Publicado: (2024)
Which LLMs are Difficult to Detect? A Detailed Analysis of Potential Factors Contributing to Difficulties in LLM Text Detection
por: Thorat, Shantanu, et al.
Publicado: (2024)
por: Thorat, Shantanu, et al.
Publicado: (2024)
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
por: Lu, Yao, et al.
Publicado: (2026)
por: Lu, Yao, et al.
Publicado: (2026)
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
por: Mukherjee, Subhojyoti, et al.
Publicado: (2025)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2025)
Offline Preference Optimization via Maximum Marginal Likelihood Estimation
por: Najafi, Saeed, et al.
Publicado: (2025)
por: Najafi, Saeed, et al.
Publicado: (2025)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
por: Mendu, Sai Krishna, et al.
Publicado: (2025)
por: Mendu, Sai Krishna, et al.
Publicado: (2025)
MALTO at SemEval-2024 Task 6: Leveraging Synthetic Data for LLM Hallucination Detection
por: Borra, Federico, et al.
Publicado: (2024)
por: Borra, Federico, et al.
Publicado: (2024)
Frustratingly Easy Task-aware Pruning for Large Language Models
por: Tian, Yuanhe, et al.
Publicado: (2025)
por: Tian, Yuanhe, et al.
Publicado: (2025)
PIE: Performance Interval Estimation for Free-Form Generation Tasks
por: Hsu, Chi-Yang, et al.
Publicado: (2025)
por: Hsu, Chi-Yang, et al.
Publicado: (2025)
SIG: Speaker Identification in Literature via Prompt-Based Generation
por: Su, Zhenlin, et al.
Publicado: (2023)
por: Su, Zhenlin, et al.
Publicado: (2023)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
por: Mu, Yongyu, et al.
Publicado: (2026)
por: Mu, Yongyu, et al.
Publicado: (2026)
FLARE: Task-agnostic embedding model evaluation through a normalization process
por: Jiang, Jingzhou, et al.
Publicado: (2026)
por: Jiang, Jingzhou, et al.
Publicado: (2026)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
por: Huang, Tiansheng, et al.
Publicado: (2024)
por: Huang, Tiansheng, et al.
Publicado: (2024)
HarmPot: An Annotation Framework for Evaluating Offline Harm Potential of Social Media Text
por: Kumar, Ritesh, et al.
Publicado: (2024)
por: Kumar, Ritesh, et al.
Publicado: (2024)
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
por: Pan, Zhuoshi, et al.
Publicado: (2024)
por: Pan, Zhuoshi, et al.
Publicado: (2024)
Ejemplares similares
-
LegalLens Shared Task 2024: Legal Violation Identification in Unstructured Text
por: Hagag, Ben, et al.
Publicado: (2024) -
Investigating and Alleviating Harm Amplification in LLM Interactions
por: Guo, Ruohao, et al.
Publicado: (2026) -
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
por: Qi, Jingyuan, et al.
Publicado: (2023) -
Data Contamination Report from the 2024 CONDA Shared Task
por: Sainz, Oscar, et al.
Publicado: (2024) -
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
por: Chehbouni, Khaoula, et al.
Publicado: (2024)