Week 13 — Lesson 3: Latency and resource monitoring

3 min

A technician who waits three seconds for every answer stops using the NorthPeak agent. To make it faster you first need to know where the seconds go: in Python, in the network, in reading the prompt, or in writing the answer. Measure the wall time with time.perf_counter, read the token counts from Ollama's response, and you know how fast your model is and where the time is spent. This lesson covers the two clocks, tokens per second, and what ollama ps says about memory and context.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Week 13 — Observability and model reliability” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one