0

Xử lý OpenAI API Rate Limits trong Production: Retry, Queue và Backpressure

Một lời gọi OpenAI API thành công trong môi trường development chưa cho biết hệ thống sẽ phản ứng thế nào khi hàng trăm request đến cùng lúc, token usage tăng đột biến hoặc upstream trả về 429.

Cách xử lý phổ biến nhất là thêm retry:

request failed → sleep → retry

Cách này có thể hoạt động khi traffic thấp. Nhưng khi nhiều process cùng retry, chính retry lại tạo thêm tải, chiếm connection, kéo dài queue và làm rate-limit incident nghiêm trọng hơn.

Một pipeline production cần nhiều lớp kiểm soát hơn:

Request → Admission control → Queue hoặc execution slot → OpenAI API → Bounded retry → Result / degraded response / dead-letter queue

Bài viết này tập trung vào implementation mechanics của pipeline đó. Các con số về concurrency, retry và timeout trong ví dụ chỉ là cấu hình minh họa; mỗi hệ thống phải chọn chúng từ SLO, workload và giới hạn thực tế của project.

OpenAI API Production Request Flow


1. Why retry-only logic fails


Giả sử 100 request đồng thời nhận 429 và tất cả đều retry sau đúng hai giây. Hai giây sau, hệ thống tạo ra một burst mới gồm 100 request.

Nếu burst thứ hai tiếp tục bị giới hạn, vòng lặp sẽ trở thành:

429 → đồng loạt chờ → đồng loạt retry → 429

Đây là một dạng retry storm.

Retry-only còn có bốn vấn đề khác:

  • Request đang chờ retry vẫn chiếm tài nguyên của application.
  • Một tenant có thể sử dụng hết concurrency của các tenant còn lại.
  • Request đã hết giá trị với người dùng vẫn tiếp tục chạy.
  • Nếu SDK và application đều retry, số lần gọi thực tế có thể lớn hơn dự kiến.

OpenAI lưu ý rằng request không thành công vẫn có thể tính vào giới hạn theo phút. Vì vậy, retry liên tục không giúp hệ thống thoát khỏi rate limit.

Nguyên tắc đầu tiên là:

Retry phải có giới hạn về số lần, tổng thời gian và lượng tài nguyên được phép tiêu thụ.


2. Classifying API failures correctly


Không phải lỗi nào cũng nên retry.

Một error handler production nên xét ít nhất:

  • HTTP status
  • error.type
  • error.code
  • Retry-After
  • deadline còn lại của request
  • request đã tạo side effect hay chưa

Bảng phân loại khởi đầu:

Temporary throttling - 429 với lỗi rate limit tạm thời - Tôn trọng Retry-After, sau đó retry có giới hạn. Temporary overload - 503 hoặc server overload - Retry có backoff nếu còn deadline. Network interruption - Connection reset trước khi nhận response - Retry nếu operation an toàn. Timeout - Attempt vượt timeout - Chỉ retry nếu tổng deadline vẫn còn. Authentication - 401 - Không retry; sửa credential/configuration. Authorization - 403 - Không retry tự động. Invalid request - 400 - Không retry cùng payload. Billing/quota requiring action - Lỗi cần thay đổi billing hoặc limit - Không retry tự động. Safety/policy rejection - Request không được chấp nhận - Không biến thành retry loop.

Theo tài liệu OpenAI hiện tại, Retry-After có thể xuất hiện trên 429 do rate limit tạm thời và 503 do model overload. Điều đó không có nghĩa các lỗi quota, billing hoặc lỗi cần hành động của người vận hành cũng có thể được giải quyết bằng retry.

Một classifier đơn giản:

type Decision =
  | { action: "retry"; reason: string }
  | { action: "defer"; reason: string }
  | { action: "fail"; reason: string };

function classifyFailure(error: ApiError): Decision {
  const operatorActionCodes = new Set([
    "organization_usage_limit_exceeded",
    "organization_spend_limit_exceeded",
    "project_spend_limit_exceeded",
    "credit_balance_exhausted",
    // Defensive handling for older or compatibility responses.
    "insufficient_quota",
  ]);

  // A 429 can mean temporary throttling or one of several limits that require
  // billing, credit, or usage configuration changes. Never retry the latter.
  if (
    typeof error.code === "string" &&
    operatorActionCodes.has(error.code)
  ) {
    return { action: "fail", reason: "operator_action_required" };
  }

  if (error.status === 429 && error.code === "slow_down") {
    return { action: "retry", reason: "temporary_rate_limit" };
  }

  if (error.status === 429 && error.type === "rate_limit_error") {
    return { action: "retry", reason: "temporary_rate_limit" };
  }

  if (error.status === 503 && error.code === "server_is_overloaded") {
    return { action: "retry", reason: "temporary_model_overload" };
  }

  if (error.status === 401 || error.status === 403) {
    return { action: "fail", reason: "credentials_or_permissions" };
  }

  if (error.status === 400) {
    return { action: "fail", reason: "invalid_request" };
  }

  if (error.isNetworkError) {
    return { action: "retry", reason: "network_failure" };
  }

  return { action: "fail", reason: "unclassified_failure" };
}

Các error code có thể thay đổi theo endpoint và SDK. Đừng chỉ dựa vào status code nếu response body cung cấp thông tin chi tiết hơn.


3. Bounded exponential backoff with jitter


Exponential backoff tăng thời gian chờ sau mỗi lần thất bại:

base → base × 2 → base × 4 → base × 8 → ...

Nếu mọi worker dùng cùng một lịch chờ, chúng vẫn có thể thức dậy cùng lúc. Jitter thêm một khoảng ngẫu nhiên để phân tán retry.

Một retry wrapper nên có:

  • giới hạn số attempt
  • giới hạn tổng thời gian
  • delay tối đa
  • jitter
  • support cho Retry-After
  • cancellation
  • retry budget dùng chung với SDK
import OpenAI from "openai";

const client = new OpenAI({ maxRetries: 0 });

type RetryOptions = {
  maxAttempts: number;
  maxElapsedMs: number;
  baseDelayMs: number;
  maxDelayMs: number;
};

const sleep = (ms: number, signal?: AbortSignal) =>
  new Promise<void>((resolve, reject) => {
    const timer = setTimeout(resolve, ms);
    signal?.addEventListener("abort", () => {
      clearTimeout(timer);
      reject(new Error("request_cancelled"));
    }, { once: true });
  });

async function createResponseWithRetry(
  input: string,
  options: RetryOptions,
  signal?: AbortSignal,
) {
  const startedAt = Date.now();
  let attempt = 0;

  while (attempt < options.maxAttempts) {
    attempt += 1;

    try {
      return await client.responses.create({
        model: process.env.OPENAI_MODEL!,
        input,
      });
    } catch (error: any) {
      const decision = classifyFailure(normalizeError(error));
      if (decision.action !== "retry") throw error;

      const elapsedMs = Date.now() - startedAt;
      const remainingMs = options.maxElapsedMs - elapsedMs;

      if (attempt >= options.maxAttempts || remainingMs <= 0) {
        throw new Error("retry_budget_exhausted", { cause: error });
      }

      const retryAfterMs = parseRetryAfter(error);
      const exponentialMs = options.baseDelayMs * 2 ** (attempt - 1);
      const cappedMs = Math.min(exponentialMs, options.maxDelayMs);
      const fallbackDelayMs = Math.random() * cappedMs;

      const delayMs = retryAfterMs !== null
        ? retryAfterMs + Math.random() * options.baseDelayMs
        : fallbackDelayMs;

      if (delayMs > remainingMs) {
        throw new Error("deadline_would_be_exceeded", { cause: error });
      }

      await sleep(delayMs, signal);
    }
  }

  throw new Error("retry_budget_exhausted");
}

Điểm quan trọng không nằm ở một công thức backoff duy nhất. Hệ thống cần bảo đảm:

attempt_count <= configured_limit elapsed_time <= operation_deadline SDK retries + application retries <= total_retry_budget

Các OpenAI SDK chính thức có thể tự retry một số lỗi đủ điều kiện. Nếu application thêm retry bên ngoài, cần tắt SDK retry hoặc tính nó vào cùng một budget để tránh nested retry.


4. Respecting Retry-After


Khi OpenAI cung cấp Retry-After, application nên xem đây là thời gian chờ tối thiểu.

Không nên cap Retry-After xuống local maximum rồi retry sớm hơn yêu cầu. Nếu server delay lớn hơn deadline còn lại, chuyển request sang queue, trả trạng thái pending hoặc fail có kiểm soát.

function parseRetryAfter(error: any): number | null {
  const raw = error?.headers?.get?.("retry-after")
    ?? error?.headers?.["retry-after"];

  if (!raw) return null;

  const seconds = Number(raw);
  if (Number.isFinite(seconds) && seconds >= 0) {
    return seconds * 1000;
  }

  const dateMs = Date.parse(raw);
  if (Number.isFinite(dateMs)) {
    return Math.max(0, dateMs - Date.now());
  }

  return null;
}

Ngoài Retry-After, response có thể cung cấp header về request limit, token limit, remaining capacity và reset time. Các header này phù hợp cho telemetry và adaptive admission control, nhưng application không nên giả định chúng luôn xuất hiện.


5. When to introduce a queue


Synchronous execution phù hợp khi người dùng đang chờ trực tiếp, tác vụ ngắn, deadline chặt và lưu lượng tương đối ổn định.

Queue phù hợp khi workload có burst, tác vụ có thể hoàn thành bất đồng bộ, cần retry sau một khoảng dài, cần fairness giữa tenant hoặc cần dead-letter queue.

type OpenAIJob = {
  jobId: string;
  tenantId: string;
  operation: string;
  payloadRef: string;
  idempotencyKey: string;
  createdAt: string;
  deadlineAt: string;
  attempt: number;
  maxAttempts: number;
  priority: "interactive" | "standard" | "batch";
};

Không nên đẩy toàn bộ dữ liệu nhạy cảm vào message nếu queue không được thiết kế để lưu dữ liệu đó. Có thể lưu payload ở data store phù hợp rồi đưa payloadRef vào queue.

async function processJob(job: OpenAIJob) {
  if (Date.now() >= Date.parse(job.deadlineAt)) return markExpired(job);
  if (await resultExists(job.idempotencyKey)) return markDuplicate(job);

  const admission = await limiter.acquire({
    tenantId: job.tenantId,
    estimatedTokens: await estimateTokens(job),
  });

  if (!admission.allowed) return reschedule(job, admission.retryAt);

  try {
    const result = await executeWithRetryBudget(job);
    await saveResultOnce({ idempotencyKey: job.idempotencyKey, result });
    await markCompleted(job);
  } catch (error) {
    const decision = classifyFailure(normalizeError(error));
    if (decision.action === "retry" && job.attempt + 1 < job.maxAttempts) {
      return reschedule(job, nextAvailableTime(error));
    }
    return moveToDeadLetterQueue(job, error);
  } finally {
    await admission.release();
  }
}

Queue không tự giải quyết overload. Nếu producer tạo job nhanh hơn consumer xử lý trong thời gian dài, backlog vẫn tăng vô hạn. Vì vậy queue phải đi cùng admission control và backpressure.

Với workload không cần phản hồi tức thời, Batch API cũng là một lựa chọn cần đánh giá. Batch API uses a separate rate-limit pool from standard per-model limits. Batch là luồng asynchronous; batch có thể hết hạn và request chưa hoàn thành sẽ bị hủy.


6. Concurrency and per-tenant admission control


Rate limit upstream và concurrency nội bộ là hai vấn đề khác nhau. Ngay cả khi chưa chạm RPM, application vẫn có thể gặp quá nhiều connection mở, memory tăng, worker starvation hoặc tenant lớn chiếm hết execution slot.

class Semaphore {
  private active = 0;
  private readonly waiters: Array<() => void> = [];

  constructor(private readonly limit: number) {}

  async acquire(): Promise<() => void> {
    if (this.active >= this.limit) {
      await new Promise<void>((resolve) => this.waiters.push(resolve));
    }

    this.active += 1;
    return () => {
      this.active -= 1;
      this.waiters.shift()?.();
    };
  }
}

const openAISemaphore = new Semaphore(
  Number(process.env.OPENAI_MAX_CONCURRENCY),
);

Global semaphore vẫn chưa bảo đảm fairness. Nên có thêm:

global concurrency AND tenant concurrency AND tenant request budget AND tenant token/cost budget

Không hard-code OpenAI limit vào source code. Limits có thể phụ thuộc model, project, usage tier và cấu hình account.


7. Backpressure patterns


Backpressure là cách hệ thống báo cho producer rằng consumer hiện không thể tiếp nhận thêm công việc với tốc độ hiện tại.

Không có backpressure:

traffic tăng → queue tăng → latency tăng → request hết hạn → retry tăng → queue tăng thêm

Các pattern thực tế:

  • Reject early: Nếu request không thể hoàn thành trong SLO, trả lỗi có kiểm soát ngay tại admission layer.
  • Bounded queue: Giới hạn số job, tổng kích thước, tuổi job, backlog theo tenant và priority class.
  • Load shedding: Loại bỏ hoặc trì hoãn workload ít quan trọng trước.
  • Producer throttling: Giảm ingest khi queue depth vượt threshold.
  • Deadline propagation: Truyền deadline từ client qua API, queue và worker.
  • Circuit breaker: Tạm ngừng phần lớn request mới khi tỷ lệ lỗi upstream tăng cao và dùng probe request để kiểm tra recovery.

8. Prompt/request caching considerations


Application response cache phù hợp với extraction từ nội dung bất biến, deterministic classification hoặc request lặp lại với cùng input và policy.

Cache key nên gồm:

tenant model configuration prompt version tool/schema version normalized input hash authorization scope

Không cache khi output cần real-time data, request có side effect, permission context có thể thay đổi hoặc hai tenant không được chia sẻ kết quả.

Request coalescing giảm duplicate in-flight calls:

const inFlight = new Map<string, Promise<unknown>>();

async function coalesce<T>(key: string, operation: () => Promise<T>): Promise<T> {
  const existing = inFlight.get(key);
  if (existing) return existing as Promise<T>;

  const current = operation().finally(() => inFlight.delete(key));
  inFlight.set(key, current);
  return current;
}

OpenAI prompt caching tái sử dụng phần prefix giống nhau trên request đủ điều kiện. Cache hit không được bảo đảm. Để tăng reuse, đặt instruction/reference ổn định trước, nội dung thay đổi sau, giữ tool definition ổn định và theo dõi cached-token usage.

Caveat quan trọng: cached input tokens vẫn được tính vào tokens-per-minute limits. Prompt caching có thể giúp latency và input cost nhưng không phải cách bỏ qua rate limit.


9. Observability and cost telemetry


Mỗi operation nên có correlation ID và telemetry ở ba lớp.

Request-level: request_id, trace_id, tenant_id, user_id_hash, operation, priority, prompt_version, model, queue_wait_ms, attempt_count, retry_delay_ms, total_latency_ms, outcome, error_status, error_code.

Usage: input_tokens, output_tokens, cached_input_tokens, cache_write_tokens_when_available, estimated_cost, usage_source.

Rate control: concurrency_in_use, tenant_concurrency_in_use, queue_depth, oldest_job_age_ms, retry_budget_remaining, rate-limit remaining/reset fields.

async function recordUsage(response: any, context: RequestContext) {
  const usage = response.usage;

  await telemetry.write({
    requestId: context.requestId,
    tenantId: context.tenantId,
    operation: context.operation,
    model: response.model,
    inputTokens: usage?.input_tokens,
    outputTokens: usage?.output_tokens,
    cachedInputTokens: usage?.input_tokens_details?.cached_tokens,
    observedAt: new Date().toISOString(),
  });
}

Không nên ghi toàn bộ prompt/output vào log mặc định. Cần xác định PII, secret, retention, access control, tenant isolation, masking và sampling.


10. Failure modes and degraded operation


Upstream temporary throttling:

  • giảm admission rate
  • tôn trọng Retry-After
  • queue workload có thể trì hoãn
  • reject sớm request sắp hết deadline

Model overload:

  • bounded retry
  • circuit breaker
  • fallback chỉ khi fallback model đã được đánh giá

Queue unavailable:

  • interactive request có thể dùng synchronous path giới hạn
  • background request fail closed hoặc ghi vào durable fallback store
  • không dùng memory queue như durability mechanism duy nhất

Worker crash after API success có thể tạo duplicate call khi queue redelivery. Cần idempotency ở application layer:

async function saveResultOnce(input: {
  idempotencyKey: string;
  result: unknown;
}) {
  return database.transaction(async (tx) => {
    const existing = await tx.results.find(input.idempotencyKey);
    if (existing) return existing;

    return tx.results.insert({
      idempotencyKey: input.idempotencyKey,
      result: input.result,
    });
  });
}

Nếu streaming bị gián đoạn sau khi client đã nhận một phần output, không tự động replay toàn bộ request rồi ghép hai stream. Hãy đánh dấu incomplete, cho phép retry thủ công hoặc tạo operation mới phù hợp.


11. Production checklist


Failure classification

  • [ ] Phân biệt temporary throttling với quota/billing/configuration error
  • [ ] Kiểm tra cả status và error code
  • [ ] Không retry authentication, authorization hoặc invalid request
  • [ ] Có xử lý riêng cho interrupted stream

Retry

  • [ ] Giới hạn số attempt
  • [ ] Giới hạn tổng retry time
  • [ ] Có exponential backoff và jitter
  • [ ] Tôn trọng Retry-After
  • [ ] SDK retry và application retry dùng chung một budget
  • [ ] Request hết deadline không tiếp tục retry

Queue và backpressure

  • [ ] Queue có giới hạn
  • [ ] Job có deadline/expiry
  • [ ] Có dead-letter queue
  • [ ] Producer được throttle khi backlog tăng
  • [ ] Có load shedding theo priority
  • [ ] Worker concurrency có thể điều chỉnh độc lập

Multi-tenant control

  • [ ] Có global concurrency limit
  • [ ] Có per-tenant concurrency limit
  • [ ] Có request/token/cost budget theo tenant
  • [ ] Một tenant không chiếm toàn bộ execution pool
  • [ ] Cache và coalescing key giữ tenant isolation

Duplicate prevention

  • [ ] Mỗi logical operation có idempotency key
  • [ ] Result được ghi theo cơ chế write-once
  • [ ] Queue redelivery không tạo duplicate side effect
  • [ ] Streaming retry không trộn hai response

Telemetry

  • [ ] Theo dõi queue wait và execution latency riêng
  • [ ] Theo dõi attempt count và retry delay
  • [ ] Lưu usage theo tenant và operation
  • [ ] Theo dõi cached token usage khi có
  • [ ] Có alert cho retry amplification và queue age
  • [ ] Không log secret hoặc nội dung nhạy cảm mặc định

12. Further reading


Official OpenAI references:

For the broader production architecture, see:

https://titanbases.com/blog/scaling-openai-api-production?utm_source=viblo&utm_medium=referral&utm_campaign=openai_enterprise_vn_2026q4&utm_content=viblo_openai_scale


All rights reserved

Viblo
Hãy đăng ký một tài khoản Viblo để nhận được nhiều bài viết thú vị hơn.
Đăng kí