An Apache Spark-based analytics platform optimized for Azure.
Hi @Andres de la Garza
The error you’re seeing is due to the token throughput limit (tokens per minute) configured for the model at the workspace level. This typically occurs when requests are either large (high token count) or multiple requests are sent in parallel.
In the meantime, you can try the following to reduce the impact:
- Reduce the input/prompt size where possible
- Limit parallel requests (avoid bursts)
- Add small delays between requests
- Implement retry with backoff for transient failures
Tools like Cursor may appear to handle this better because they internally queue or split requests, whereas direct usage can hit limits more quickly.
On the quota side, we’ll check this internally with the Databricks team to see if the current limit can be reviewed or adjusted for your workload.
To help us with that, could you please share:
- Approximate tokens per request
- Number of parallel requests being sent
- Frequency (requests per minute)
This will help us validate and proceed further.