What our V100 + vLLM stack actually feels like (Qwen2.5-7B)

Four Tesla V100s, one vLLM engine and 360 real ShareGPT conversations. Every metric defined before it is reported, every serving flag explained, and an honest read on which workloads this hardware fits.

How to read LLM inference numbers (before you trust a tok/s chart)

A big tok/s chart is not enough. Learn TTFT, ITL, goodput, and why ISL, OSL, and concurrency belong on every inference label before you trust the number.

You Probably Don’t Need AWS (And That’s Okay)

I’m going to say something that might sound controversial in tech circles: most companies using AWS, Azure, or GCP are massively over-provisioned and under-utilizing what they’re paying for. There. I said it. Before the replies…

Our Servers Heat Greenhouses. No, Seriously.

Every time someone tells me they work in “green tech,” I brace myself for the usual pitch. Carbon offsets. Renewable energy certificates. A tree planted somewhere for every signup. Nice gestures, sure. But let’s be…

Why Your Data Belongs in Europe (And Why It Matters More Than You Think)

Last year, a friend of mine — a founder running a mid-size SaaS platform — got a letter from a client in Germany. Not a happy one. The client’s legal team had flagged that user…