Jumbo, side by side
Jumbo adds three blocks of layers from Qwen3.6-27B to Qwen3.8-27B, at the depths where the weights changed most. I wanted to see whether that would recover correct answers the newer model no longer gave. The table reports the factual evaluation, including wrong answers and abstentions. You can also send your own prompt to both models at the same 8-bit quantization, with greedy decoding and thinking mode off.
Results on 2,000 factual questions
Loading the bank…
Try a prompt
Both models generate at once on one machine; up to 400 tokens each.
How to read it
The shaded text marks where the responses first differ. Each model selects its most likely next token, and tokens per second are shown separately. The evaluation above used 2,000 questions from the long tail of PopQA, with the same short-answer prompt for all three models. A recovery means Jumbo answered correctly when the parent did not; a regression means the reverse. The scoring rule accepts any response containing an approved answer alias, including some responses a reader would reject. The live demo uses a general assistant prompt, so a live response can differ from the benchmark response. See the write-up for the full comparison and its limitations.