Qdrant 官方最新动态:Oxidizing Cross-Encoders

ADK Qdrant官方 / ADK编译 2026-09-25 3分钟 72 次浏览
速览导读 / Summary

We rewrote cross-encoder inference in Rust, benchmarked it against Python, and on a small reranking model it came out only 1.1x faster at the median. That is not a sign of a slow Rust implementation.

官方发布 2026-09-25 官网即时同步

Key Insights / 核心看点

  • 1 We rewrote cross-encoder inference in Rust, benchmarked it against Python, and o

Qdrant 官方最新动态:Oxidizing Cross-Encoders

来源:Qdrant 官方动态 | 发布日期:2026-09-25

核心更新概览

We rewrote cross-encoder inference in Rust, benchmarked it against Python, and on a small reranking model it came out only 1.1x faster at the median. That is not a sign of a slow Rust implementation.

详细内容记录

We rewrote cross-encoder inference in Rust, benchmarked it against Python, and on a small reranking model it came out only 1.1x faster at the median. That is not a sign of a slow Rust implementation. The Python libraries we compared against, , hand the same work to the same C++ engine, , so all three spend almost all of their time in the same compiled code. In the Rust community, rewriting code in Rust is called “oxidizing” it, since rust is also what iron turns into when it oxidizes. This post covers how we oxidized cross-encoder inference into a library, , how it works, where it does pull ahead (about 1.4x on a larger model), and where it does not (model load time). Before we dive into the implementation, two concepts will come up frequently: : A model that reads a query and a document together and returns one relevance score for the pair. Cross-encoders are typically used for reranking: a fast first stage retrieves a pool of candidates, and the cross-encoder re-sorts them so the best matches come first. An embedding model (a bi-encoder) turns the query and each document into separate vectors and compares them afterward. A cross-encoder feeds both texts into the model at once, so every word of the query is compared with every word of the document. That makes it more accurate, and also more expensive: nothing can be computed ahead of time, so every query-document pair costs a full model run. : ONNX (Open Neural Network Exchange) is a file format for trained models. You export a model from PyTorch or TensorFlow once, and run it anywhere. ONNX Runtime is the C++ engine that runs these files, with graph optimizations, hardware acceleration, and bindings for many languages, including Rust. We used the

更多技术细节可访问官方原文:https://qdrant.tech/blog/oxidizing-cross-encoders/。