首页 > AI前沿 > Benchmarking Pocket-Scale Inference

Benchmarking Pocket-Scale Inference

Hacker News 2026-08-28 03:12 7 阅读 查看原文
Benchmarking pocket-scale inference We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process. “Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page. Intelligence and Inference Performance Summary Average Score (16K max context) vs. End-to-End Generation Time|iPhone 17 Pro Average Score (Mobile Device Benchmark Set) A simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details. End-to-End Generation Time Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns. Inference Performance End-to-End Generation Time|iPhone 17 Pro End-to-End Generation Time Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns. Model Intelligence Average Score (Mobile Device Benchmark Set, 16K max context)|iPhone 17 Pro Average Score (Mobile Device Benchmark Set) A simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details. Looking for the Artificial Analysis Intelligence Index scores for these models? The following models have been evaluated on our full index, and their scores are visible on their model pages: Gemma 4 31B (Reasoning) LFM2.5-2.6B Ling 3.0 Tiny Ministral 3 3B Ministral 3 8B Qwen3.5 9B (Reasoning) Qwen3.6 27B (Reasoning) Qwen3.6 35B A3B (Reasoning) Token Efficiency Context Budget Overruns Evaluation Breakdown Mobile Device Benchmark Set Evaluations (16K max context)|iPhone 17 Pro BFCL Tool calling (index subset) IFBench Instruction following AA-Omniscience Accuracy Knowledge AA-Omniscience Non-Hallucination Rate 1 - hallucination rate GPQA Diamond Scientific reasoning MATH-500 Quantitative reasoning