CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Guoheng, Wang, Ziyao, Tian, Bowei, Liu, Meng, Shen, Zheyu, He, Shwai, He, Yexiao, Ye, Wanghao, Wang, Yiting, Li, Ang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services
by: Sun, Guoheng, et al.
Published: (2025)
by: Sun, Guoheng, et al.
Published: (2025)
Fair Diagnosis: Leveraging Causal Modeling to Mitigate Medical Bias
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025)
by: Shen, Zheyu, et al.
Published: (2025)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
MindCraft: How Concept Trees Take Shape In Deep Models
by: Tian, Bowei, et al.
Published: (2025)
by: Tian, Bowei, et al.
Published: (2025)
Towards counterfactual fairness through auxiliary variables
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Prada: Black-Box LLM Adaptation with Private Data on Resource-Constrained Devices
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
Revisiting Federated Fine-Tuning: A Single Communication Round is Enough for Foundation Models
by: Wang, Ziyao, et al.
Published: (2024)
by: Wang, Ziyao, et al.
Published: (2024)
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
by: Wang, Ziyao, et al.
Published: (2024)
by: Wang, Ziyao, et al.
Published: (2024)
CogniPair: From LLM Chatbots to Conscious AI Agents -- GNWT-Based Multi-Agent Digital Twins for Social Pairing -- Dating & Hiring Applications
by: Ye, Wanghao, et al.
Published: (2025)
by: Ye, Wanghao, et al.
Published: (2025)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
MCP4EDA: LLM-Powered Model Context Protocol RTL-to-GDSII Automation with Backend Aware Synthesis Optimization
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
Cognibit: From Digital Exhaustion to Real-World Connection Through Gamified Territory Control and LLM-Powered Twin Networking
by: Ye, Wanghao, et al.
Published: (2026)
by: Ye, Wanghao, et al.
Published: (2026)
EvoVerilog: Large Langugage Model Assisted Evolution of Verilog Code
by: Guo, Ping, et al.
Published: (2025)
by: Guo, Ping, et al.
Published: (2025)
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
by: He, Yexiao, et al.
Published: (2024)
by: He, Yexiao, et al.
Published: (2024)
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
by: Sun, Guoheng, et al.
Published: (2026)
by: Sun, Guoheng, et al.
Published: (2026)
Router-Tuning: A Simple and Effective Approach for Enabling Dynamic-Depth in Transformers
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Demystifying When Pruning Works via Representation Hierarchies
by: He, Shwai, et al.
Published: (2026)
by: He, Shwai, et al.
Published: (2026)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
by: He, Yexiao, et al.
Published: (2025)
by: He, Yexiao, et al.
Published: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
by: Yu, Hanfei, et al.
Published: (2026)
by: Yu, Hanfei, et al.
Published: (2026)
Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Understanding and Harnessing Sparsity in Unified Multimodal Models
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
NeuroSymAD: A Neuro-Symbolic Framework for Interpretable Alzheimer's Disease Diagnosis
by: He, Yexiao, et al.
Published: (2025)
by: He, Yexiao, et al.
Published: (2025)
Making Large Language Models Efficient Dense Retrievers
by: Lei, Yibin, et al.
Published: (2025)
by: Lei, Yibin, et al.
Published: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
Nanomolding single-crystalline CoIn3 and RhIn3 nanowires
by: Duong, Nghiep Khoan, et al.
Published: (2025)
by: Duong, Nghiep Khoan, et al.
Published: (2025)
Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL
by: Yao, Zhewei, et al.
Published: (2025)
by: Yao, Zhewei, et al.
Published: (2025)
SCARA: A Semantics-Constrained Autonomous Remediation Agent for Opaque Industrial Software Vulnerabilities
by: Ning, Bowei, et al.
Published: (2026)
by: Ning, Bowei, et al.
Published: (2026)
Towards Building Non-Fine-Tunable Foundation Models
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Token-Efficient Change Detection in LLM APIs
by: Chauvin, Timothée, et al.
Published: (2026)
by: Chauvin, Timothée, et al.
Published: (2026)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
Beyong Tokens: Item-aware Attention for LLM-based Recommendation
by: Zhang, Xiaokun, et al.
Published: (2026)
by: Zhang, Xiaokun, et al.
Published: (2026)
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
Similar Items
-
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
by: Wang, Ziyao, et al.
Published: (2025) -
Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services
by: Sun, Guoheng, et al.
Published: (2025) -
Fair Diagnosis: Leveraging Causal Modeling to Mitigate Medical Bias
by: Tian, Bowei, et al.
Published: (2024) -
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025) -
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)