My Favorite (and Least Favorite) LLMs

· fizzy blog


(last updated: July 26th, 2026)

In no particular order (well except for the first one which is by far my favorite of the current breed of models), these are the LLMs I personally use.

My Usecases #

The additional context that my personal tasks do not primarily revolve around agentic coding, math, and in general I mostly do not use LLMs for STEM tasks. Most of my usage revolves around more creative tasks (i.e. textual analysis and synthesis, as I enjoy reading other takes on media I enjoy and LLMs can actually have interesting takes, or at least interestingly worded takes), as well as some conversational research tasks using the Kagi MCP.

I do sometimes use LLMs for conversational coding though, like for example (a real prompt I asked):

how would i write a PF rule to allow all in on a specific interface wt0?

When I use LLMs in assistant interfaces I mostly use a variation of the Claude system prompt which has been changed and altered using Anthropic-like introspection from models (most of the alteration work came from Claude Fable 5 when it was available, but much came from MiMo as well).

The Good #

MiMo v2.5 (non-pro and pro) #

I do not know how Xiaomi of all companies did it, but MiMo is genuinely one of the best generalized models (and best modern open weights one) outside of STEM tasks, in my experience. It writes creatively and is able to write thought-out analysis, it has good introspective reasoning abilities (as is evidenced by being the other model I used for system prompt re-writing), it doesn't overthink when unnecessary, it's just... a really good model. The Pro version is somewhat smarter but the non-pro version is very, very good for its size and price as well.

It's also the best model I know of for "we have Claude at home," it has very similar neuroses (especially if you tell it its Claude). Not necessarily the "you're absolutely right" (it does do that sometimes to be fair) but things like when it decides to use lists, when it asks clarifying questions, how it hedges its answers, how it adds extra sections like this: instead of with markdown headers or such.

Ironically it's not quite as good at STEM tasks but it's like. Good enough that I can use it as a first pass for those types of questions without worrying too much about its answers, and if they seem suspicious I can get a second 'opinion' from another model.

Update July 26th: Yep, still great :)

Kimi K3 #

So remember all the things I said about Kimi the first time? That it was smart and all but the personality left a little bit to be desired? Well, K3 has knocked everything else out of the park, including GPT 5.5 for me, in smarts, but with a slightly more bearable personality and better prompt following! Due to its comical size it also has insanely good world knowledge

Deepseek v4 Pro and Flash #

Very solid runner-ups though I wouldn't call it especially close overall. However, Deepseek is very (relatively) STEMmaxxed, making it good for a second smarter opinion for those types of questions. Flash is more creative when asked to write or analyze, but Pro is smarter for STEM questions; however they're both really quite close at the end of the day.

Update July 26th: I think with K3 existing there's definitely less of a place for the Deepseeks now. Hopefully they come out of preview swinging again :3

Kimi K2 0711 #

The original K2 is still very good if you want super creative takes on something, even if its not as smart as more modern models. Providers are starting to drop support though so it probably won't last that long.

Good but Local #

Admittedly, I can't actually run many of these locally, but these ones are the ones in the range where a lot of people can run them locally. There are, of course, only really two options here:

Gemma 4 #

It's good. Like, it's not the second coming of Satan, but it's pretty good, especially for being a local 20-30b model range. The small models (E2B and E4B) are also very good for what they are!

Qwen 3.5/3.6/3.7 #

They're good for agent and STEM stuff, I guess. They also finetune easier than Gemma 4 and have more available sizes, though nothing is quite as good as E2B and E4B in the lower range.

The Mediocre #

GLM 5/5.1/5.2 #

This might be a hot take, but I haven't liked any Zhipu models since maybe GLM 4.7 Flash. They haven't been smarter enough to justify their continued lurch towards being blander and blander, and even the messianic GLM 5.2 has not modified this trend, in my testing.

Update July 26th: And now it's been outdone by K3, all while K3 is much better generalized :D

Claude Opus 4.5+ #

Don't get me wrong, they're pretty smart, but as time goes on they've been getting dumber and dumber in practice for me. Whether that's in search of Claude Code-maxxing by Anthropic, I don't know. Their style has also been getting worse and worse. Opus 4.5 was good in its day, but I would argue open models have sufficiently surpassed it in all fields that matter to me.

Claude Fable 5 #

It very occasionally has interesting vibes but it's so insanely expensive you'd be better off running 100 K3s in parallel. Admittedly I did use it for introspecting on my system prompt and got some nice bonuses from that, so... eh.

GPT 5.6 Sol #

Very big model vibes (similar to K3 a bit) but it needs to reason so much and so expensively for a good response that, again, in terms of cost you're best off running like 4 K3s in parallel.

The Really Bad #

All Ernies #

"Who is Ernie?" Don't ask questions you don't want to know the answer to.

All Lings #

They aren't awful base models but dear god they're simultaneously so unstable and so bland at the same time, and have nothing going for them vs literally everything else.

last updated: