This is better than GLM-5.2 which just 2 months ago was my go to high tier model!
Also in OpenRouter this model is listed to have a maximum context of 1.31M tokens! this is the first high-end model to my knowledge that supports more than 1M tokens context 🤯
So this was the Ox-Alpha model on OpenRouter. That’s cool.
I guess the sleuths were right. I’ve used local GLM models in the past, they aren’t bad at all. although I’m currently enjoying the speed and quality of Ornith 1.5 35b running on HaloFPX.
just wish there was an abliterated version.
There are several abliterated versions listed in the finetunes: https://huggingface.co/models?other=base_model%3Afinetune%3Aornith-ai%2FOrnith-1.5-35B-A3B
I haven’t tested any of them personally though.
oh I know, but since I’m running on a Z13, there’s a custom runtime built for Strix Halo that makes the model a bit faster. and that runtime uses a fine-tuned version of Ornith that takes advantages of the specific architecture to increase t/s. I know, it’s already fast, but it’s nice to see those first tokens start generating within a couple seconds, and my agent crank out a few thousand tokens in less than ten seconds.
ornith isn’t good for coding, sure, but for automating the boring shit I don’t want to pay attention to, the model works beyond expectations. I do wish it was abliterated because one of the things I use it for is OSINT workflows and tracking down the personal info of fash, and the guardrails version doesn’t always like to do that, so the skills have to carefully word the prompts. I imagine it would get pissy if I asked it hard questions about China too, but I don’t chat with LLMs so I’m not sure.




