What is an NPU?
The chip your laptop and phone now advertise — what it's for, and why its headline "TOPS" number tells you less than it looks.
The chip your laptop and phone now advertise — what it's for, and why its headline "TOPS" number tells you less than it looks.
Teaching a model how to behave — not what is true.
Asking for the working, and getting a better answer as a side effect.
It is slow, and it also silently gave your model a 4K context window. Both are the same bug.
It is not the model. Ollama picks your context length from your VRAM at startup, and the bottom tier is very small.
The file format behind every local model you have downloaded — and how to read its name.
A model that can take actions in a loop — and the word the industry has worn out.
The current frontier and open-weight language models, with every spec read from a primary source.
Turning meaning into coordinates — the trick that makes semantic search work.
The 2017 architecture underneath essentially every model you have heard of.
When a model states something false with total confidence — and why that is the normal case, not a glitch.
Paying once for the part of your prompt that never changes — and the write fee nobody mentions.
Giving the model the documents instead of hoping it memorised them.
The model's working memory — and the reason a long chat starts forgetting things.
Four levers, ranked by how much they actually save. Switching vendor is not one of them.
Shrinking a model so it fits on hardware you own — and what you give up.
They are the same engine. You are choosing an interface and a licence, not a speed.
Output costs five to six times more than input — at every vendor, at every tier. That ratio, not the headline price, is what decides your bill.
Google gives away every Flash text model and an older Pro. It does not give away the current Pro, or a single image.