Exercise 2 — The nearest incidents with embeddings

Guided practice5 min
Time
20-30 min
You need
the kit, the venv, python data/make_dataset.py done, Ollama running with nomic-embed-text pulled
Deliverable
the three nearest incidents of INC-0001 (ids and scores), pasted in a text file

The lab kit of the course: https://github.com/hrhouma2/aiopsatlas-ml-data-diagnostics-labs-en

Goal

A technician opens incident INC-0001 and asks: "Did we see this before?" You will answer with embeddings. Every description becomes a vector of 768 numbers from nomic-embed-text. A cosine similarity between vectors says how close two incidents are. You will find the three nearest incidents of INC-0001, then search the file with a sentence written in your own words.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Week 8 — NLP and text representations” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one