A new quantized version of the 27-billion-parameter Qwen3.8 model shrinks the original 54.7 GB checkpoint to 15.7 GB while adding only 0.80% perplexity. The OrcaSAQ2 27B Cyber Uncensored GGUF, released by OrcaRouter, is built for local deployment on consumer GPUs and targets coding, tool use, and authorized security testing workloads.
The quantization uses a proprietary sensitivity-aware mixed-precision system. On a 24 GB GPU, the model runs at 20.5 tokens per second with full offload, or 27.6 tokens per second when paired with DFlash2 speculative decoding. The compressed checkpoint achieves 94.4% token-level Top-1 agreement with the original BF16 reference and a mean KLD of 0.020.
What fits on a single GPU
The 15.7 GB footprint means the model can run entirely in VRAM on hardware that cannot hold the original 54.7 GB checkpoint. On a 24 GB card, full offload with DFlash2 enabled peaks at 18.1 GB - leaving headroom for context and batching. A practical starting point is roughly 32K interactive context with DFlash2 on, though the architecture supports up to 262K tokens.
On a 16 GB GPU, the model still runs at full offload without a drafter, delivering 20.5 tok/s. DFlash2 speculative decoding adds 35% single-stream throughput by drafting blocks of tokens in one pass and verifying them losslessly against the model.
Uncensored and text-only
This checkpoint is derived from an abliterated version of Qwen3.8-27B with the refusal direction removed. It will attempt requests that a safety-tuned model would decline. The release documentation states that guardrails, filtering, and policy enforcement are the deployer's responsibility. The model is text-only - the vision tower is not included.
Quantization is not mathematically lossless. The 94.4% Top-1 agreement means some token decisions differ from the BF16 reference. OrcaRouter notes that the +0.80% perplexity measurement reflects model fidelity and does not guarantee identical downstream task performance. Detailed methodology for the quantization system has not been disclosed.
Deployment and serving
The checkpoint runs in llama.cpp, Ollama, and LM Studio. An OpenAI-compatible API is available through llama-server. Recommended sampling parameters are temperature 1.0, top_p 0.95, and top_k 20. Thinking mode is enabled by default. For agent workloads, OrcaRouter advises benchmarking against the actual tool schema and context distribution used in production.
DFlash2 drafters for this base model are available publicly at Q8_0 precision, weighing 2.06 GB. The quantization inherits the Apache-2.0 license from the original Qwen3.8-27B model.
Why this matters for developers and security engineers
The practical result is a 27B-class reasoning model that fits on a single consumer GPU without cloud costs. For red teams, vulnerability researchers, and developers building coding agents or tool-heavy applications, that means uncensored behavior and 262K context in a local deployment envelope. The trade-off is a small fidelity gap versus the full-precision reference - measurable at 0.80% perplexity and 94.4% token agreement - which teams should validate against their own eval suites before committing to production pipelines.
Your membership also unlocks: