Skip to content

< all problems06 · Level 02, Search

Generate Embeddings with Hugging Face

medium · implement · Embeddings & Retrieval

Implement embed(sentences, tokenizer, model), returning a (n, d) NumPy array (not a tensor) of normalized sentence embeddings.

  1. Tokenize the batch with padding=True, truncation=True, return_tensors="pt".
  2. Run the model under torch.no_grad().
  3. Masked mean pooling. The model returns one vector per token. Average them into one per sentence, over real tokens only.
  4. L2-normalize each row.

The catch: padding positions carry attention mask value 0. Average them in and every short sentence drifts toward the same meaningless point.