GRPO pilot: correcting geopolitical bias in model weights
By Daniel Melo Avila
A judge-guided GRPO study shows adapter-only training can raise Taiwan-domain neutrality from 0.479 to 0.967 on a 35B MoE model—without full fine-tuning.
Prompt-only debiasing barely moves strong geopolitical preferences in large language models. This pilot asks a sharper question: can domain-specific reinforcement learning correct that bias directly in the weights?
The study freezes a Qwen3.6-35B-A3B backbone, trains a compact LoRA adapter with QLoRA, and optimizes with group-relative policy optimization (GRPO) using judge-only rewards from an external Llama 3.3 70B evaluator. Training used 887 Taiwan-explicit prompts under a no-system-prompt protocol so measured gains can be attributed to the adapter rather than prompt engineering.
On a held-out suite of 60 Taiwan prompts, the step-350 adapter raised mean judge neutrality from 0.479 to 0.967—closing 93.7% of the achievable neutrality gap while preserving general capabilities on fixed canary probes. The full paper below includes the methods, equations, training curves, and evaluation tables.
Full paper















More from Ainslie
Choosing where Ainslie runs AI
Why Ainslie One combines on-box models with managed cloud capacity instead of forcing every task down the same path.
Read article →What private AI means in Ainslie One
A clear look at the boundary between company data on Ainslie One and the hosted services that keep access simple.
Read article →Why company-built apps need a home
Useful AI-built tools should remain available to the team, improve over time, and become part of everyday work.
Read article →