Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
Fuente:
arXiv
Salvato in:
| Autori principali: | Ackerman, Christopher, Panickssery, Nina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
Mitigating Many-Shot Jailbreaking
di: Ackerman, Christopher M., et al.
Pubblicazione: (2025)
di: Ackerman, Christopher M., et al.
Pubblicazione: (2025)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
di: Ball, Sarah, et al.
Pubblicazione: (2024)
di: Ball, Sarah, et al.
Pubblicazione: (2024)
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
di: Ackerman, Christopher
Pubblicazione: (2026)
di: Ackerman, Christopher
Pubblicazione: (2026)
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2024)
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2024)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
di: Yu, Erxin, et al.
Pubblicazione: (2025)
di: Yu, Erxin, et al.
Pubblicazione: (2025)
Assessing the Emergent Symbolic Reasoning Abilities of Llama Large Language Models
di: Petruzzellis, Flavio, et al.
Pubblicazione: (2024)
di: Petruzzellis, Flavio, et al.
Pubblicazione: (2024)
Refusal in Language Models Is Mediated by a Single Direction
di: Arditi, Andy, et al.
Pubblicazione: (2024)
di: Arditi, Andy, et al.
Pubblicazione: (2024)
AgentInstruct: Toward Generative Teaching with Agentic Flows
di: Mitra, Arindam, et al.
Pubblicazione: (2024)
di: Mitra, Arindam, et al.
Pubblicazione: (2024)
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
di: Baek, In-Chang, et al.
Pubblicazione: (2025)
di: Baek, In-Chang, et al.
Pubblicazione: (2025)
Llama-Nemotron: Efficient Reasoning Models
di: Bercovich, Akhiad, et al.
Pubblicazione: (2025)
di: Bercovich, Akhiad, et al.
Pubblicazione: (2025)
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
di: Li, Xiaochuan, et al.
Pubblicazione: (2024)
di: Li, Xiaochuan, et al.
Pubblicazione: (2024)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
di: Crispino, Nicholas, et al.
Pubblicazione: (2023)
di: Crispino, Nicholas, et al.
Pubblicazione: (2023)
Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
di: Zheng, Haoyang, et al.
Pubblicazione: (2025)
di: Zheng, Haoyang, et al.
Pubblicazione: (2025)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
BanglaLlama: LLaMA for Bangla Language
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
Open Llama2 Model for the Lithuanian Language
di: Nakvosas, Artūras, et al.
Pubblicazione: (2024)
di: Nakvosas, Artūras, et al.
Pubblicazione: (2024)
Self-Evolving Critique Abilities in Large Language Models
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
di: Martra, Pere
Pubblicazione: (2025)
di: Martra, Pere
Pubblicazione: (2025)
Controlled Generation for Private Synthetic Text
di: Zhao, Zihao, et al.
Pubblicazione: (2025)
di: Zhao, Zihao, et al.
Pubblicazione: (2025)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
di: Yoon, Junsang, et al.
Pubblicazione: (2024)
di: Yoon, Junsang, et al.
Pubblicazione: (2024)
Lugha-Llama: Adapting Large Language Models for African Languages
di: Buzaaba, Happy, et al.
Pubblicazione: (2025)
di: Buzaaba, Happy, et al.
Pubblicazione: (2025)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
di: Schröder, Christopher, et al.
Pubblicazione: (2024)
di: Schröder, Christopher, et al.
Pubblicazione: (2024)
Automated Text Scoring in the Age of Generative AI for the GPU-poor
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio
di: Xi, Ningyuan, et al.
Pubblicazione: (2024)
di: Xi, Ningyuan, et al.
Pubblicazione: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
di: Liu, Wanlong, et al.
Pubblicazione: (2024)
di: Liu, Wanlong, et al.
Pubblicazione: (2024)
Badllama 3: removing safety finetuning from Llama 3 in minutes
di: Volkov, Dmitrii
Pubblicazione: (2024)
di: Volkov, Dmitrii
Pubblicazione: (2024)
Self-Recognition in Language Models
di: Davidson, Tim R., et al.
Pubblicazione: (2024)
di: Davidson, Tim R., et al.
Pubblicazione: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
di: Hu, Yujia, et al.
Pubblicazione: (2024)
di: Hu, Yujia, et al.
Pubblicazione: (2024)
Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy
di: Cullhed, Eric
Pubblicazione: (2024)
di: Cullhed, Eric
Pubblicazione: (2024)
Scalability Matters: Overcoming Challenges in InstructGLM with Similarity-Degree-Based Sampling
di: Lee, Hyun, et al.
Pubblicazione: (2025)
di: Lee, Hyun, et al.
Pubblicazione: (2025)
Text2Data: Low-Resource Data Generation with Textual Control
di: Wang, Shiyu, et al.
Pubblicazione: (2024)
di: Wang, Shiyu, et al.
Pubblicazione: (2024)
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
di: Ali, Mehdi, et al.
Pubblicazione: (2024)
di: Ali, Mehdi, et al.
Pubblicazione: (2024)
RKadiyala at SemEval-2024 Task 8: Black-Box Word-Level Text Boundary Detection in Partially Machine Generated Texts
di: Kadiyala, Ram Mohan Rao
Pubblicazione: (2024)
di: Kadiyala, Ram Mohan Rao
Pubblicazione: (2024)
On the Ability of Transformers to Verify Plans
di: Sarrof, Yash, et al.
Pubblicazione: (2026)
di: Sarrof, Yash, et al.
Pubblicazione: (2026)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
di: Ajwani, Rohan Deepak, et al.
Pubblicazione: (2024)
di: Ajwani, Rohan Deepak, et al.
Pubblicazione: (2024)
A Regularization-based Transfer Learning Method for Information Extraction via Instructed Graph Decoder
di: Chen, Kedi, et al.
Pubblicazione: (2024)
di: Chen, Kedi, et al.
Pubblicazione: (2024)
Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque
di: Corral, Ander, et al.
Pubblicazione: (2024)
di: Corral, Ander, et al.
Pubblicazione: (2024)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
di: Toshniwal, Shubham, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023) -
Mitigating Many-Shot Jailbreaking
di: Ackerman, Christopher M., et al.
Pubblicazione: (2025) -
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
di: Ball, Sarah, et al.
Pubblicazione: (2024) -
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
di: Ackerman, Christopher
Pubblicazione: (2026) -
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2024)