Skill · Backend
Upstash ratelimit
Add rate limiting to API routes, middleware, and edge functions with @upstash/ratelimit: sliding window, fixed window, and token bucket backed by Upstash Redis.
How to use it
- Start your plan and connect your AI once
- Ask for the task in your own words, or say it directly:
Use the Upstash ratelimit skill to help me with this.Without a connection: copy the SKILL.md below into your AI's project instructions.
Upstash Ratelimit
Overview
@upstash/ratelimit implements distributed rate limiting on top of Upstash Redis. Because state lives in Redis, every instance of a serverless function or edge worker shares the same counters, which an in-memory limiter cannot do. It ships three algorithms (fixed window, sliding window, token bucket), per-identifier keys, optional in-memory blocking of already-limited identifiers, and optional analytics.
When to Use This Skill
- Use when the user needs to limit requests per IP, user, API key, or tenant
- Use when protecting login, signup, form, webhook, or LLM endpoints from
- Use when choosing between fixed window, sliding window, and token bucket.
- Do not use for client-side retry/backoff against a third-party API's limits;
- Do not use for a single long-running process with no shared state; an
across multiple serverless instances or regions.
abuse and returning 429 Too Many Requests.
see api-rate-limit-handler.
in-memory limiter is simpler there.
How It Works
Step 1: Install and configure
npm install @upstash/ratelimit @upstash/redis
Set UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN in the environment.
Step 2: Create the limiter once, outside the handler
import { Ratelimit } from "@upstash/ratelimit";
import { Redis } from "@upstash/redis";
export const ratelimit = new Ratelimit({
redis: Redis.fromEnv(),
limiter: Ratelimit.slidingWindow(10, "10 s"), // 10 requests per 10 seconds
prefix: "rl:api",
analytics: true,
});
Constructing the limiter at module scope lets the built-in ephemeral cache short-circuit blocked identifiers without a Redis call.
Step 3: Call limit() with a stable identifier
const { success, limit, remaining, reset, pending } = await ratelimit.limit(userId);
success is false when the identifier is over its limit. reset is a Unix timestamp in milliseconds. pending is a promise for background work (analytics, multi-region sync); await it or pass it to waitUntil on edge runtimes so the function is not frozen before it completes.
Examples
Example 1: Next.js middleware returning 429
import { Ratelimit } from "@upstash/ratelimit";
import { Redis } from "@upstash/redis";
import { NextResponse, type NextRequest } from "next/server";
const ratelimit = new Ratelimit({
redis: Redis.fromEnv(),
limiter: Ratelimit.slidingWindow(20, "1 m"),
});
export async function middleware(request: NextRequest) {
const ip = request.headers.get("x-forwarded-for") ?? "anonymous";
const { success, limit, remaining, reset } = await ratelimit.limit(ip);
if (!success) {
return new NextResponse("Too Many Requests", {
status: 429,
headers: {
"X-RateLimit-Limit": String(limit),
"X-RateLimit-Remaining": String(remaining),
"X-RateLimit-Reset": String(reset),
"Retry-After": String(Math.ceil((reset - Date.now()) / 1000)),
},
});
}
return NextResponse.next();
}
export const config = { matcher: "/api/:path*" };
Example 2: Token bucket with per-plan limits
const limiters = {
free: new Ratelimit({
redis: Redis.fromEnv(),
prefix: "rl:free",
limiter: Ratelimit.tokenBucket(5, "10 s", 10), // refill 5 per 10 s, burst 10
}),
pro: new Ratelimit({
redis: Redis.fromEnv(),
prefix: "rl:pro",
limiter: Ratelimit.tokenBucket(50, "10 s", 100),
}),
};
const { success } = await limiters[plan].limit(apiKey);
Best Practices
- ✅ Use a stable, low-cardinality identifier (user id, API key, tenant) where
- ✅ Set a distinct
prefixper endpoint or plan so limits do not collide. - ✅ Return
Retry-AfterandX-RateLimit-*headers with 429 responses. - ✅ Prefer
slidingWindowfor most APIs; usetokenBucketwhen short bursts - ❌ Don't construct a new
Ratelimitinside the request handler. - ❌ Don't rely on
pendingcompleting on its own in edge runtimes. - ❌ Don't rate limit by
x-forwarded-forwithout validating it is set by
possible; fall back to IP only for anonymous traffic.
are acceptable; use fixedWindow when the lowest Redis cost matters.
your proxy; clients can spoof it otherwise.
Limitations
- Requires an Upstash Redis database; it does not work with other Redis
- Each
limit()call is at least one HTTP round trip to Redis, so it adds - Sliding window is an approximation that assumes an even spread of requests
MultiRegionRatelimittrades strict accuracy for lower latency and does- If Redis is unreachable, the default
timeout(5 s) lets requests through - This skill does not replace environment-specific validation, testing, or
servers or without network access.
latency to every request it guards.
in the previous window; it is not an exact log.
not support the token bucket algorithm.
(reason: "timeout"); this fails open, not closed.
expert review.
Security & Safety Notes
- Rate limiting is one layer of abuse protection, not authentication. Pair it
- The Redis token grants full database access; keep it server-side.
- Changing limits in production can lock out legitimate users. Confirm the
with auth and input validation.
numbers with the user before deploying stricter limits.
Common Pitfalls
- Problem: Every request is allowed even after the limit.
- Problem: Analytics are empty on Vercel Edge or Cloudflare Workers.
Solution: Each identifier must be the same string across requests; check that the identifier is not undefined or a fresh random value.
Solution: Pass pending to waitUntil (ctx.waitUntil(pending)) so the background request is not cancelled when the response is sent.
Related Skills
@upstash-redis- The client this package uses for storage.@api-rate-limit-handler- Client-side backoff and retry when you are the@upstash-qstash- Queue and smooth traffic to downstream services instead
one being rate limited.
of rejecting it.