PipelineScore
← Back to leaderboard
local

unsloth/gpt-oss-20b-GGUF

Released Context 0Kunsloth-gpt-oss-20b-gguf
PipelineScore
0.0DRIP
Ranked #15 of 20 models · 25th percentileThroughput is the headline (50.9); RAG is the soft spot (0.0). Best-fit profile: Local-first.

Category breakdown

Score per category, normalized 0–100 against the v1 anchor.

Code
0.0
Reason
0.0
Tool Use
0.0
RAG
0.0
Speed
50.9

Strengths

Speed50.9
Code0.0
Reason0.0

Same model, different rigs

Every submission of unsloth/gpt-oss-20b-GGUF on the 0–100 scale. The spread is the point: where it runs changes what you get.

0255075100

Best 79.6 on i7-6850k-cpu-32gb · lowest 0.0 on i7-6850k-cpu-32gb · spread 79.6 pts across 5 runs. Hover a dot for its rig.

Sample tasks

A taste of what the test pack measures. Full prompts are private and rotated daily.

CodeDifficulty 1code-fib-1

Fibonacci function

Write a Python `fib(n)` returning the nth Fibonacci number, O(n).

ReasonDifficulty 1reason-math-1

Train meeting time

Two trains, opposite directions, given speeds and start times — when do they meet?

RAGDifficulty 2rag-extract-1

Extract metrics to JSON

From the context, extract net sales, operating margin, and free cash flow as a JSON object. Numbers only.

Tool UseDifficulty 2tool-schema-1

OpenAPI param selection

Given an OpenAPI schema with limit/offset/sort, fill JSON for 'next 50, recent first.'

RAGDifficulty 2rag-grounding-1

Refuses to fabricate

Context lacks the answer — does the model fabricate or correctly say it can't?

Compare with

unsloth/gpt-oss-20b-GGUF · PipelineScore