Skip to content
AI News HubLIVE
Original source1 min read

EmbeddingGemma 2 and Why Apache 2.0 Matters for Embedding Models

Summary

In a comment on the Hacker News thread about EmbeddingGemma 2, Simon Willison welcomes the model's Apache 2.0 license and argues that embedding models in particular should not be closed, proprietary and hosted-only: once a vendor retires a model, users must pay to re-embed the millions of vectors they have already stored. He does not want to self-host, but wants the open weights as a fallback if a provider ever stops serving the model.

EmbeddingGemma 2 and Why Apache 2.0 Matters for Embedding Models
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

Comment: EmbeddingGemma 2

Simon Willison’s Weblog

Subscribe

6th October 2026

Comment My comment on EmbeddingGemma 2 — Hacker News

I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.

For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.

Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.

If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.

(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)

Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.

Recent articles

We're going to need default hard budget caps on pretty much everything - 3rd October 2026

OpenAI DevDay 2026 live blog - 29th September 2026

2026 in LLMs (so far) - 27th September 2026

This is a beat by Simon Willison, posted on 6th October 2026.

google 417

ai 2,264

generative-ai 2,007

embeddings 62

gemma 17

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe

Disclosures

Colophon

©

2002

2003

2004

2005

2006

2007

2008

2009

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

2021

2022

2023

2024

2025

2026

Key points and analysis

Article intelligence

EngineersIntermediate

Key points

  • Simon Willison praises EmbeddingGemma 2 for shipping under the Apache 2.0 license, arguing embedding models especially should not be closed, hosted-only and proprietary.
  • Embedding workflows involve computing and storing thousands or millions of vectors, so a vendor retiring a model forces users to pay to recompute all existing embeddings.
  • He notes OpenAI offered in April 2024 to cover the financial cost of users re-embedding content with new models, but says that cannot be relied on from every provider.
  • He prefers paying a provider to host the model rather than hosting it himself, while knowing he could run the open weights or move to another vendor if hosting stops.

Highlights and analysis are generated automatically and may contain errors. Check the original source.