Scaling from 8B to 14B Yields No Meaningful Improvement in Biomimetic Prompt Following: A Paired Comparison Across 3 Model Families and 35 Configurations
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | COYAUD, Denis |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prompt Engineering: a methodology for optimizing interactions with AI-Language Models in the field of engineering
von: Juan David Velásquez-Henao
Veröffentlicht: (2023)
von: Juan David Velásquez-Henao
Veröffentlicht: (2023)
The Invisible Chaperone: The Secret World of System Prompts
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 598: Audio Engineering as Prompt Engineering: Better Sound, Better AI
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 1086: Why AI Can't Stop Talking About Second Order Effects
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
The Transformer Trinity: Why Three Architectures Rule AI
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 1111: The Architecture of Intelligence: Beyond the Transformer
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 651: Decoding the Blueprint: An Expert Guide to AI Model Cards
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 111: Beyond Transformers: Solving the AI Memory Crisis
von: Rosehill, Daniel, et al.
Veröffentlicht: (2025)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2025)
Ep. 1080: Beyond the Prompt: Mapping the Future of Claude Opus
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 103: The Future of Coding: Is Your Brain Wired for AI?
von: Rosehill, Daniel, et al.
Veröffentlicht: (2025)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2025)
The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models
von: Dubey, Mradul
Veröffentlicht: (2026)
von: Dubey, Mradul
Veröffentlicht: (2026)
Ep. 170: The Heavy Metal of Machine Learning: Inside PyTorch
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Why AI Can't Simulate Extreme Decision-Making
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
von: Khare, Mohit
Veröffentlicht: (2026)
von: Khare, Mohit
Veröffentlicht: (2026)
Ep. 23: AI's Blind Spot: Data, Bias & Common Crawl
von: Rosehill, Daniel, et al.
Veröffentlicht: (2025)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2025)
Ep. 713: The AI Cyber Frontier: Israel as a Global Testing Ground
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Ep. 476: Beyond the Plateau: AI-Powered Language Mastery in 2026
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Perturbing LLM Attractors, Intentionally: The Thermodynamics of Human-AI Interaction
von: Pourdavood, Parham
Veröffentlicht: (2025)
von: Pourdavood, Parham
Veröffentlicht: (2025)
When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings
von: Kolb, Christian
Veröffentlicht: (2026)
von: Kolb, Christian
Veröffentlicht: (2026)
Local Large Language Models in R with Ollama
von: Schweinberger, Martin
Veröffentlicht: (2026)
von: Schweinberger, Martin
Veröffentlicht: (2026)
Ep. 1108: Beyond the Emoji: How Hugging Face Conquered AI
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Reliability Inference Drives Cue Extraction in Large Language Models Consuming External Reasoning Traces
von: HIDEKI
Veröffentlicht: (2026)
von: HIDEKI
Veröffentlicht: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Tuning PC v3
von: Abhinav Gorantla
Veröffentlicht: (2026)
von: Abhinav Gorantla
Veröffentlicht: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
von: Ertugrul Coban
Veröffentlicht: (2025)
von: Ertugrul Coban
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context CB-StaticDiscovery v1
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Benchmark: VAR-LiNGAM, PCMCIplus v3
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
von: Pratanu Mandal
Veröffentlicht: (2026)
von: Pratanu Mandal
Veröffentlicht: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
von: Pratanu Mandal
Veröffentlicht: (2025)
von: Pratanu Mandal
Veröffentlicht: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
von: Pratanu Mandal
Veröffentlicht: (2026)
von: Pratanu Mandal
Veröffentlicht: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v3
von: Ertugrul Coban
Veröffentlicht: (2025)
von: Ertugrul Coban
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Benchmark run results by Shu Wan, on benchmark context PC Hyperparameter Tuning v2
von: Shu Wan
Veröffentlicht: (2025)
von: Shu Wan
Veröffentlicht: (2025)
Beyond Buttons: Is the Admin Dashboard Dead?
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
Ep. 869: Why Tiny Digital Savants Are Outperforming God-Models
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
The Absurdist's Guide to AI Probing: How I Learned to Stop Worrying and Love the Nonsense
von: Walton, Mathew
Veröffentlicht: (2026)
von: Walton, Mathew
Veröffentlicht: (2026)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
von: Ivković, Jovan
Veröffentlicht: (2026)
von: Ivković, Jovan
Veröffentlicht: (2026)
Design and Construction of a Snake-Like Robot Implementing Rectilinear and Sidewinding Gait Motions
von: Jairo José Marín Arciniegas
Veröffentlicht: (2023)
von: Jairo José Marín Arciniegas
Veröffentlicht: (2023)
Fortifying NLP models - dataset + code
von: Ferdinan, Teddy, et al.
Veröffentlicht: (2025)
von: Ferdinan, Teddy, et al.
Veröffentlicht: (2025)
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
Ähnliche Einträge
-
Prompt Engineering: a methodology for optimizing interactions with AI-Language Models in the field of engineering
von: Juan David Velásquez-Henao
Veröffentlicht: (2023) -
The Invisible Chaperone: The Secret World of System Prompts
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026) -
Ep. 598: Audio Engineering as Prompt Engineering: Better Sound, Better AI
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026) -
Ep. 1086: Why AI Can't Stop Talking About Second Order Effects
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026) -
The Transformer Trinity: Why Three Architectures Rule AI
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)