Skip to main content

Rate Limits

caution

The Weave API is available to customers on the Weave Enterprise plan only

caution

Note: Figma reserves the right to change rate limits. Changes may affect specific endpoints or plans.

Figma applies rate limits to requests to the Weave API to provide a consistent and reliable experience for users. One flat limit applies to every endpoint:

API endpointsLimit
All Weave API endpoints150/min

Rate limits are tracked per token, per endpoint. Every request to the Weave API is authenticated with a Weave API token, and each token gets its own budget for each endpoint. Two consequences follow:

  • Each endpoint is budgeted separately. Spending a token's full allowance on POST runs does not rate-limit that token's calls to GET runs.
  • Each token is budgeted separately. A workspace can have multiple active tokens, and one token being rate-limited does not affect the others.

The budget is per endpoint, not per tool: runs of every tool in the workspace are counted against the same POST runs budget.

How rate limiting works​

For managing rate limits, Figma uses a leaky bucket algorithm. A token's bucket for an endpoint holds 150 requests and refills continuously at 150 per minute, so a burst of up to 150 requests is allowed and sustained traffic is held to the per-minute rate. When the Weave API is unable to fulfill requests due to the bucket being full, the endpoint returns a 429 error.

429 errors

A 429 returns the following response body:

{
"status": 429,
"err": "Rate limit exceeded"
}

The response carries the following header:

"Retry-After": Integer

The following table describes the header in the error.

FieldTypeDescription
Retry-AfterIntegerIn seconds, how long before you should retry sending the request.

What if my integration is hitting rate limits?​

Runs are asynchronous, so the request pattern that hits the limit is usually a polling loop, not the runs themselves. If your integration is getting rate-limited:

  • Replace polling with webhooks. Registering a webhook when you start a run means Weave calls you as each run reaches a terminal state, which removes the polling loop entirely.
  • If you do poll, poll in batches and back off. GET runs takes a comma-separated run_ids list, so one request can cover every run you're waiting on. Runs take a variable amount of time, so poll on an interval that reflects how long your tool usually takes rather than as fast as the limit allows.
  • Cache what doesn't change. A published tool version's GET inspect response and the GET tools list only change when the workspace's tools do. Fetch them once and reuse them across runs instead of re-fetching them for every request.
  • Handle 429 correctly. Make sure your integration retries after the Retry-After value rather than immediately.