EPISODE · Jun 18, 2026 · 3 MIN
@liquidai:Liquid AI 發布了 LFM2.5-Embedding-350M 與 LFM2.5-ColBERT-350M,提供 11 種語言的超高速與高精確度檢索能力…
from EasyVibeCoding Podcast · host Liquid AI
Liquid AI 發布了 LFM2.5-Embedding-350M 與 LFM2.5-ColBERT-350M,提供 11 種語言的超高速與高精確度檢索能力。 這兩款 350M 參數的模型是 Liquid AI 旗下「LFM」家族首批雙向編碼器,專為短文本情境(如產品目錄、FAQ、知識庫)設計,並針對企業級應用優化,端到端檢索延遲最低可達 1.5ms。 模型定位與差異 Liquid AI 指出,檢索模型通常需要在速度與準確度之間做出取捨,這兩款模型分別對應不同的需求: LFM2.5-Embedding-350M:採用密集雙編碼器架構,將每個文件轉換為單一向量。適合追求極致搜尋速度與最小索引體積的場景。 LFM2.5-ColBERT-350M:採用「Late Interaction」架構,將每個 token 轉換為向量,並透過 MaxSim 進行逐詞比對。適合對準確度要求較高、可容忍較大索引體積的場景。 Liquid AI 推出的 LFM2.5-ColBERT-350M 與 LFM2.5-Embedding-350M 在多語言檢索基準測試(NanoBEIR 與 MKQA)中,表現皆優於其前代版本及其他同類模型(如 Qwen3-Embedding-0.6B、gte-multilingual-base 等),展現出同級最佳的檢索品質。 技術架構與效能 兩款模型均基於 LFM2.5-350M-Base 進行雙向補丁(bidirectional patches)調整,將原本的因果解碼器(causal decoder)轉變為雙向編碼器,使每個 token 都能同時關注左右上下文。 部署靈活性:支援 llama.cpp 的 GGUF 格式,可在 CPU、筆記型電腦及邊緣裝置上運行,端到端查詢嵌入延遲低於 10ms。 企業級效能:在 GPU 上透過自定義執行環境(custom runtime),端到端查詢嵌入延遲可低於 2ms。 語言支援:涵蓋阿拉伯語、德語、英語、西班牙語、法語、義大利語、日語、韓語、挪威語、葡萄牙語及瑞典語。 在單張 H100 GPU 的內部 GPU 堆疊測試中,LFM2.5-ColBERT Query Embedding 在高併發數(Concurrency=32)下展現出最高的每秒查詢率(QPS),突破 5,000 QPS,顯著優於加入 MaxSim 或 Doc Embedding 的檢索配置。 實作指引(以 LFM2.5-ColBERT-350M 為例) 若要使用 PyLate 進行文件索引與檢索,請依照下列步驟操作: 安裝必要套件: `bash pip install -U pylate ` 載入模型並初始化索引: `python from pylate import indexes, models, retrieve 載入模型 model = models.ColBERT( modelnameor_path="LiquidAI/LFM2.5-ColBERT-350M", trustremotecode=True, ) model.tokenizer.padtoken = model.tokenizer.eostoken 初始化 PLAID 索引 index = indexes.PLAID( index_folder="pylate-index", index_name="index", override=True, ) ` 編碼並新增文件: `python documents = ["document 1 text", "document 2 text", "document 3 text"] documents_ids = ["1", "2", "3"] documents_embeddings = model.encode( documents, batch_size=32, is_query=False, showprogressbar=True, ) index.add_documents( documentsids=documentsids, documentsembeddings=documentsembeddings, ) ` 執行檢索: `python retriever = retrieve.ColBERT(index=index) queries_embeddings = model.encode( ["query for document 3", "query for document 1"], batch_size=32, is_query=True, showprogressbar=True, ) scores = retriever.retrieve( queriesembeddings=queriesembeddings, k=10, ) ` 更多詳細資訊可參考 官方部落格文章、Hugging Face 頁面 或 官方文件。原文:https://easyvibecoding.app/curated/2097
Embed this episode
Ready to play
@liquidai:Liquid AI 發布了 LFM2.5-Embedding-350M 與 LFM2.5-ColBERT-350M,提供 11 種語言的超高速與高精確度檢索能力…
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.