How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

Imported from official source

Research

AI Classified by Officially

Of course, making AI research accessible requires a powerful search engine, so that humans and agents can quickly find relevant and related work, either through the website or the pwc search CLI command, which agents can use via the Skill.

It's important to note that searching for research is not quite the same as searching for regular text. A useful paper search engine should find an exact title or arXiv identifier, but it should also understand a query such as “small language models for code generation” even when those words do not appear together in a paper. It needs to recognize that “the original BERT paper” is a navigational request, tolerate an incomplete title or typos, and still respond quickly when a model service is cold or temporarily unavailable.

For Papers with Code, we built this as a hybrid search system. This is also based on our prior experience at ML6, where we developed RAG-based systems for clients. It turned out that hybrid search typically outperforms keyword- and vector-based search systems, as it combines the best of both worlds (see also this blog for more info). Keyword search finds exact mentions, whereas vector search finds more fuzzy, semantically similar terms. Note that rerankers (also called cross-encoders) can further improve the results, although they also come with additional overhead and latency.

Papers with Code relies on a PostgreSQL database, hence its full-text search capabilities provide a fast lexical baseline. For dense embeddings, pgvector is used to add semantic recall, and the reciprocal rank fusion (RRF) algorithm combines the two. Three Hugging Face services are used for the dense embeddings:

  • Hugging Face Jobs gives us burstable GPU compute for embedding the paper corpus.
  • Hugging Face Storage Buckets provides the durable handoff between our database, experiments, and Jobs.
  • Hugging Face Inference Endpoints serves low-latency embeddings for live queries and incremental updates.
  • This is an extract. The publication continues at the source.

    Read the original at the source: https://huggingface.co/blog/pwc-search

    Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Hugging Face — imported from official source
    Official source
    https://huggingface.co/blog/feed.xml RSS
    Imported
    September 15, 2026 19:08
    Versions
    1 recorded
    Identity
    https://huggingface.co/blog/pwc-search

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.