Cito

Cito is a literature retrieval engine that searches academic papers for AI research agents. It provides a JSON API and native MCP endpoint so developers can build deep research tools without hitting rate limits.

Cito

About Cito

Cito is a hybrid search engine that indexes the Semantic Scholar corpus of 236 million academic papers. It combines BM25 keyword search with SPECTER2 dense vector retrieval, fused through Reciprocal Rank Fusion (RRF) and reranked by a cross-encoder. The tool returns ranked papers with abstracts, citation counts, open-access PDF links, and DOIs—and it's built specifically as a retrieval engine for AI agents, not as a chatbot.

Review

Cito launched this week to solve a specific frustration: AI agents doing deep literature research get throttled by academic APIs. Semantic Scholar defaults to 1 request per second, which stalls agents that fire dozens of queries in a single pass. Cito indexes the corpus independently and serves results from a single CPU box in under half a second, with a free API tier set at 100 requests per minute.

Key Features

  • Hybrid search across 236M papers using BM25 keyword matching and SPECTER2 dense vectors, fused with RRF and reranked by a cross-encoder
  • Native MCP endpoint that lets Claude Code, Cursor, and other MCP-compatible agents search the literature directly with a single command
  • Batch endpoints (/search/batch for 50 queries per request, /paper/batch for up to 1,000 paper IDs) that reduce round trips during bulk operations
  • Direct DOI and arXiv lookup for resolving specific papers without going through search
  • Free web search with no signup required, plus a plain JSON API with per-key rate limits

Pricing and Value

Cito is free. The web search requires no account. API access uses free keys with a fixed-window rate limit of 100 requests per minute per key. Paper lookups via /paper/batch are not metered; only search queries count against the limit. The maker has stated that the limit is a per-key value and can be raised for specific use cases upon request.

Pros

  • Rate limit of 100 req/min per key is a fixed-window counter, not a pacer—agents can fire all 100 requests in the first second without stalling
  • MCP integration works out of the box with Claude Code, according to early user reports
  • Batch endpoints handle up to 1,000 paper IDs in a single call, which cuts down on round trips during reference resolution
  • Returns structured metadata including citation counts, open-access PDF links, and DOIs alongside ranked results
  • Single CPU box serving sub-half-second response times keeps operational costs low, which aligns with the free pricing model

Cons

  • The 100 req/min limit is shared across all parallel workers using the same API key—ten workers don't get ten times the capacity
  • No streaming endpoint or async job pattern for long-running queries, though the maker notes that normal queries return quickly and batch endpoints cover most high-volume scenarios
  • Cito is not suited for users who want a conversational AI research assistant; it's a retrieval engine that returns ranked papers, not synthesized answers

Cito fits into workflows where an AI agent handles the reasoning layer and needs a fast, unthrottled pipeline to academic paper metadata. Researchers building custom literature review agents or running deep-research passes with Claude Code will find the MCP endpoint and batch operations immediately useful. Anyone expecting a chatbot-style interface that summarizes findings in natural language should look elsewhere—this tool hands off the thinking to your agent and sticks to retrieval.



Open 'Cito' Website
Get Daily AI Tools Updates

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)

Join thousands of clients on the #1 AI Learning Platform

Explore just a few of the organizations that trust Complete AI Training to future-proof their teams.