跳到主要內容
AI News HubLIVE
來源內容 · 翻譯待補全6 分鐘閱讀

待翻譯:miniCOIL EN-ES: Sparse Neural Retrieval Across the Language Barrier

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Sparse neural retrieval finally started getting more and more attention (SPARSEUP, MILCO, sparse encoders in Sentence Transformers…). Well, amazing, we like attention! Our miniCOIL v1 sparse neural retriever article ended with a promise: to keep working on this sparse neural retriever in-depth, improving the model’s quality, and in-width, “extending it to other dense encoders and to languages beyond English.” Here is the next “in-width” step of the miniCOIL saga: to try and cross the language barrier, that is, make the model suitable for multilingual retrieval. It’s a fun challenge, as sparse retrievers are hard to adapt to this scenario: exact matching in cross-lingual text search usually means translation.

來源Qdrant Blog作者: [email protected] (Andrey Vasnetsov)
待翻譯:miniCOIL EN-ES: Sparse Neural Retrieval Across the Language Barrier
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Learn Technical Articles miniCOIL EN-ES: Sparse Neural Retrieval Across the Language Barrier Back to Embedding Research miniCOIL EN-ES: Sparse Neural Retrieval Across the Language Barrier Juan Pablo Sotelo and Evgeniya Sukhodolskaya · October 08, 2026 Sparse neural retrieval finally started getting more and more attention (SPARSEUP, MILCO, sparse encoders in Sentence Transformers…). Well, amazing, we like attention! Our miniCOIL v1 sparse neural retriever article ended with a promise: to keep working on this sparse neural retriever in-depth, improving the model’s quality, and in-width, “extending it to other dense encoders and to languages beyond English.” Here is the next “in-width” step of the miniCOIL saga: to try and cross the language barrier, that is, make the model suitable for multilingual retrieval. It’s a fun challenge, as sparse retrievers are hard to adapt to this scenario: exact matching in cross-lingual text search usually means translation. Together with the team at Pento, we came up with miniCOIL EN-ES: a bilingual sparse neural retriever where “home” and “casa” land in the same cells of a sparse vector, cells representing one concept. This way, one inverted index serves English and Spanish queries and documents alike, no translation step in between. miniCOIL EN-ES: “home” and “casa” land in the same slot of a sparse vector. We’re sharing the design and a first set of results. We’d love your feedback: whether you’d use something of the kind, whether the approach works for you. If so, we’d scale the model’s vocabulary and training, and let’s see what comes next: the idea is easily transferable to other languages. Same Word, Different Language Word-based sparse retrieval has a limitation that no amount of relevance-based weighting alone can fix: it is based on matching exact terms: “bat” and “murciélago” can mean the same thing, but to BM25 and co. they don’t: they are two different tokens in two different postings lists. Translate, Then Search The most common approach is to translate (documents, or queries, or even both). For example, if you’d like to stick with BM25, you could run the query through a machine translation model, then do BM25-based retrieval in the corpus language. It works, but it puts a translation model in the hot path of a query, adds a dependency, and leads to management of a separate index (or a separate translated copy of the corpus) per language pair. Dense Multilingual Retrieval Multilingual dense encoders solve cross-lingual matching by construction: “She won a prize” and “Ella ganó un premio” land close together in embedding space. (At least in theory: many multilingual encoders group texts by language as well as by meaning, as our recent blog on language bias shows.) Yet dense retrieval is not a 1-1 alternative to full-text search. On the contrary, in many use cases it would shine combined with something more controllable, less recall-broad, as sometimes you’re really just asking for “I need THIS specific keyword to be in the document.” Making Sparse Vectors Polyglots We want cross-lingual matching with no translation step and no reliance on dense retrieval alone. Instead, we want to reuse the miniCOIL recipe: keep BM25’s reliable signals, a word’s importance in the document (term frequency, normalized by document length) and its rarity across the corpus (IDF), and throw on top a small learned component that adds a word’s meaning-in-the-context. However, in miniCOIL v1, the unit of the vocabulary is the English word: there’s no cell in a miniCOIL v1 sparse vector for “murciélago”, and if there were, it wouldn’t be the same cell as “bat”. So the question that defined the next edition of the miniCOIL saga, “miniCOIL EN-ES”, was: what should a sparse slot represent, if not a word? From Words to Concepts That’s how a concept of a concept was born (applause): a small set of English and Spanish words that translate to each other, such as {cat, gato} or {bat, murciélago, bate}. Concepts are not specific to English and Spanish; the approach is transferable to other languages. One Slot per Concept While in miniCOIL v1, every English word in a 30k vocabulary got its own trainable layer-model and its own 4 consecutive cells-dimensions in the sparse vector, in miniCOIL EN-ES, every concept gets one trainable head in a similar manner. When an English document says “bat” and a Spanish query says “murciélago”, both resolve to the same concept ID, both are projected by the same head, and both write into the same block of the sparse vector (in this model, 8 consecutive dimensions). That allows for both mono- and cross-lingual matching simultaneously, and still keeps miniCOIL’s ability to resolve meaning-in-the-context. One slot per concept: in miniCOIL EN-ES, “bat”, “murciélago”, and “bate” write into the same cells, and the values follow the sense; unresolved words fall back to BM25. Standing on the Shoulders of BM25, Again (We Got Comfortable) At this point, we’re jumping on these shoulders with the comfort of a clingy toddler. The scoring formula between sparse miniCOIL EN-ES vectors copy-cats the miniCOIL v1 approach. For every concept $c_i$ the query words map to, we look for documents whose words map to the same concept, in whichever language the document is written: $$ \text{score}(D,Q) = \sum_{c_i \in \text{concepts}(Q)} \text{IDF}(c_i) \cdot \text{Importance}^{c_i}_{D} \cdot {\color{YellowGreen}\text{Meaning}^{c_i(Q) \times c_i(D)}} $$ Here $c_i(Q)$ and $c_i(D)$ are the query’s and the document’s vectors for concept $c_i$; if several words of one text map to the same concept, they are pooled into a single vector before scoring. The Importance component is the BM25 term weight the concept occurrence gets (k, avg_len, and b should be adjusted according to corpus statistics on concepts; however, consider this version with usual defaults a miniCOIL-EN-ES-let’s-check-how-it-goes). The Meaning component is the dot product of two low-dimensional (eight for this model) miniCOIL vectors, as in v1. Note: The BM25 formula is suitable not only for words: the probabilistic relevance framework behind it states “the model is not restricted to terms and term frequencies — any property, attribute or feature of the document, or of the document–query pair, which we reasonably expect to provide some evidence as to relevance, may be included” (Robertson & Zaragoza, 2009, Section 2.4). For every word that isn’t in the concept vocabulary, we keep the approach from v1: fall back to plain BM25. In an ideal world, whatever isn’t part of this vocabulary is a named entity, a specific term, a code: something “untranslatable”. We hope it matches exactly, one to one, no matter the language. Of course, it’s a weak assumption, and we’d like to try other fixes for out-of-vocabulary cross-lingual matches, like the LexEcho head in MILCO; however, that’s for the next iterations. Building a Concept Vocabulary Before we can train a head per concept, we need concepts themselves, a vocabulary of them. Dictionary as a Graph To build the concept vocabulary we use the MUSE bilingual dictionaries: about 112k English-to-Spanish and a same-ish amount of Spanish-to-English pairs. Each pair is an edge in a bipartite graph, English and Spanish words on the two sides, and a concept is a tightly connected community in that graph. Running community detection (Louvain) over the aforementioned dictionary graph gives us candidate concept clusters: We keep only clusters that contain at least one word in each language. We also cap cluster size: Louvain’s resolution parameter controls how granular the communities are (higher resolution favors smaller, tighter ones), so on any cluster that grows too large we re-run Louvain with the resolution bumped up, until the concept’s word count fits under the cap. The “bajo” Problem Our first, uncapped attempt produced a few clusters of several hundred words, a little too many. The culprit is polysemy: Spanish “bajo” means low, bass, short, and under; each of those English words has its own translations, and one such word is enough to stitch dozens of weakly related terms into one giant hairball of a concept. The fix we used is not very glamorous (as anything else heuristics-, in-my-expert-opinion-it-is-ok-based): MUSE lists translations in frequency order, so we keep only the top-2 translations per source word before building the graph. Combined with a maximum cluster size of 20, this yields about 80k clean concepts, median size 2 (a typical concept is literally just a word and its translation, like {cat, gato}) and 99th percentile size 9. Concepts are Louvain communities in a bilingual dictionary graph. Polysemous words like Spanish “bajo” stitch many unrelated terms together, so we had to cap the resulting concept size. 80k Concepts Are 30 Times Too Many To verify whether the approach works, keeping 80,000 concepts and training a layer for each seemed unnecessary (for example, we could live without the {mockingjay, sinsajo} one). To trim the concept vocabulary, we streamed English and Spanish Wikipedia, split articles into sentences, and counted how many sentences contain each concept in each language, keeping only concepts with enough evidence on both sides. This leaves roughly 12,000 concepts with enough sentences for a per-concept head to learn to light up the right sense region of its concept, in whichever of the two languages the match comes in. For the released model we went even further and kept, out of the 12k, only the concepts that retrieval queries in mMARCO dev.small (a machine-translated version of the MS MARCO passage ranking dataset, which we also use for evaluation) actually use. When we rank concepts by how many query-word occurrences they catch, the top 2,398 grab 90% of everything the full 12k vocabulary would. So this first, raw, small miniCOIL EN-ES trains only 2,398 concept heads. Training miniCOIL EN-ES The v1 recipe was: take contextual word representations from an English-only dense encoder, and learn a one-layer projection into a tiny word meaning space, calibrated on how a bigger teacher encoder arranges sentences with that word. We kept the recipe and changed the ingredients (a dad joke about paella and tacos). miniCOIL EN-ES Backbone For miniCOIL EN-ES we need a multilingual input encoder that already “knows” English and Spanish, so that the trainable head focuses on the sense, not only on figuring out translation from scratch. We tested the multilingual-e5 family on parallel EN/ES sentence pairs. All sizes retrieved the parallel counterpart 100% of the time, with an average cross-lingual cosine similarity of 0.91 to 0.93. Hence, for convenience (including availability of free inference on Qdrant Cloud), we chose the smallest backbone option: intfloat/multilingual-e5-small: 384 output dimensions, 118M parameters. Each miniCOIL EN-ES head is consequently a 384 -> 8 projection: 3,072 weights (for comparison, a v1 per-word layer is 512 -> 4, 2,048 weights). Four Quadrants (not Qdrant’s), Five Points The v1 model trained on triplets: an anchor sentence containing the word, one positive sentence with this word in the same meaning and one negative sentence, where the same word has a different meaning. And the teacher encoder decides which of the two sits closer to the anchor. Keep that objective unchanged in a bilingual setting, and nothing in it forces th [truncated for AI cost control]

展開要點與分析

文章情報

投資人進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Sparse neural retrieval finally started getting more and more attention (SPARSEUP, MILCO, sparse encoders in Sentence Transformers…). Well, amazing, we like attention! Our miniCOI…

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。