Rate Limits
The Weave API is available to customers on the Weave Enterprise plan only
Note: Figma reserves the right to change rate limits. Changes may affect specific endpoints or plans.
Figma applies rate limits to requests to the Weave API to provide a consistent and reliable experience for users. One flat limit applies to every endpoint:
| API endpoints | Limit |
|---|---|
| All Weave API endpoints | 150/min |
Rate limits are tracked per token, per endpoint. Every request to the Weave API is authenticated with a Weave API token, and each token gets its own budget for each endpoint. Two consequences follow:
- Each endpoint is budgeted separately. Spending a token's full allowance on POST runs does not rate-limit that token's calls to GET runs.
- Each token is budgeted separately. A workspace can have multiple active tokens, and one token being rate-limited does not affect the others.
The budget is per endpoint, not per tool: runs of every tool in the workspace are counted against the same POST runs budget.
How rate limiting works
For managing rate limits, Figma uses a leaky bucket algorithm. A token's bucket for an endpoint holds 150 requests and refills continuously at 150 per minute, so a burst of up to 150 requests is allowed and sustained traffic is held to the per-minute rate. When the Weave API is unable to fulfill requests due to the bucket being full, the endpoint returns a 429 error.
429 errors
A 429 returns the following response body:
{
"status": 429,
"err": "Rate limit exceeded"
}
The response carries the following header:
"Retry-After": Integer
The following table describes the header in the error.
| Field | Type | Description |
|---|---|---|
Retry-After | Integer | In seconds, how long before you should retry sending the request. |
What if my integration is hitting rate limits?
Runs are asynchronous, so the request pattern that hits the limit is usually a polling loop, not the runs themselves. If your integration is getting rate-limited:
- Replace polling with webhooks. Registering a
webhookwhen you start a run means Weave calls you as each run reaches a terminal state, which removes the polling loop entirely. - If you do poll, poll in batches and back off. GET runs takes a comma-separated
run_idslist, so one request can cover every run you're waiting on. Runs take a variable amount of time, so poll on an interval that reflects how long your tool usually takes rather than as fast as the limit allows. - Cache what doesn't change. A published tool version's GET inspect response and the GET tools list only change when the workspace's tools do. Fetch them once and reuse them across runs instead of re-fetching them for every request.
- Handle
429correctly. Make sure your integration retries after theRetry-Aftervalue rather than immediately.