Luna Responses
Results & written notes
- Processing time
- —
- Documents / hour
- —
- AI cost · CAD
- —
AGILEIM
Compare processing time, throughput and estimated AI cost.
·
Results & written notes
Choices & scores
Azure allowance comparison only. Time and CAD cost use OpenAI samples; Azure Decisions support is unverified.
Planning estimate · AI cost only · excludes storage, OCR and licensing.
Same documents and API settings. Unselected local presets use 100 / 500 MiB/s; Azure uses the entered share capacity.
| Drive | Read MiB/s | Responses docs/hour | Decisions docs/hour | Responses time | Decisions time | Responses · what sets the pace | Decisions · what sets the pace |
|---|
The AI times come from WesternIM's local AgileIM test on October 8, 2026: 36 synthetic files, two passes and 72 attempts per task per endpoint. The chart extrapolates the mean application time for successful requests. It holds that latency constant when document size, token counts or concurrency change.
Both options use the same gpt-6-luna model. Model or endpoint availability depends on your account and deployment. Decisions was in public beta when this planner was prepared.
| Endpoint | Mean time/document | Successful requests | Flagged for review |
|---|
The review counts describe this synthetic test and its thresholds. They do not predict accuracy or review rates on your documents. The run estimates exclude time spent reviewing results.
The estimated rate is the slowest of reading your files, processing AI requests and any entered request/token quotas. Reads and AI calls overlap through an ideal prefetched queue. A queue keeps workers supplied but cannot exceed the slowest sustained stage. A worker that reads and waits on AI serially will be slower. At the default 1 MiB/file and 50 calls, Responses needs about 24 MiB/s; Decisions needs about 194 MiB/s. Faster storage stops improving the result once it can keep up with AI or quota limits. Azure Files provisioned tiers use capacity-based baseline recommendations; actual provisioned IOPS and throughput may be configured separately. Temporary burst performance is excluded.
Share capacity is checked for document bytes only. Allow free-space headroom. Read rates do not account for file metadata operations, other users, extraction/OCR, file-open latency or saving results. A network cap can limit every storage tier. For office-network, internal SSD and custom drives, capacity is not checked because no drive capacity was entered. Their starting read rates are illustrative assumptions, not measured specifications. Network-drive performance depends on the connection, file server and other users; SSD performance depends on the drive and workload. Replace the read rate with a representative measurement.
Hot, Cool and Transaction Optimized are HDD pricing tiers, with shared account throughput ceilings rather than dedicated per-share throughput. The assumed read rate is separate from those ceilings. IOPS measure operations, not documents.
Throughput = minimum of read MiB/s ÷ file MiB, simultaneous calls ÷ measured seconds, RPM ÷ 60, and TPM ÷ quota tokens/request ÷ 60. Each endpoint is estimated as a separate full run. “Documents processed at once” means the configured number of active requests and the sample request time set the pace; it does not establish the AI service’s maximum capacity. The optional planning ceiling removes this parallel-request constraint while retaining the entered read, request and token limits. It assumes constant latency and excludes limits left unset. Limits of 0 are unknown and excluded; faster drives or more workers cannot overcome an entered quota.
OpenAI currently publishes Build, Launch and Grow tiers. A historical “Tier 5” label does not establish the Luna limit: check your organization’s model limits, project limits and response headers. Model pools may be shared. The OpenAI examples apply the same reference cap to both endpoints as a what-if assumption; confirm Decisions limits separately.
Azure’s Tier 1–6 reference caps depend on deployment type and subscription allocation. Global and data-zone pools may be shared as subscription-level quota management rolls out. Use your deployment’s available allocation. Azure examples fill Responses only; this planner does not establish Azure Decisions availability or quota. Latency and CAD costs still use direct OpenAI data.
Quota tokens are not necessarily billed tokens. The starting estimate is input + expected output for Responses and input for Decisions. Providers may reserve output allowances or use prompt estimates; enter that reservation in Quota tokens per request. These defaults are planning assumptions, not verified endpoint accounting. Throttling, ramp-up, daily limits, retries and competing jobs can reduce actual throughput. Leave headroom and test your workload.
OpenAI rate limits · Azure quota tiers · Azure token reservations
All estimated costs are Canadian dollars. OpenAI's published billing prices are converted using an editable exchange-rate factor. The default is 1.4240, from the Bank of Canada's October 8, 2026 daily rate. It is a dated planning rate, not a live currency quote. Taxes and currency-conversion fees are excluded.
AI cost = documents × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Decisions has no output charge. Responses output includes billable reasoning. Defaults are 2,000 input and 300 Responses output tokens per document; these are illustrative, not measured benchmark usage. Adjust them in Cost settings.
One standard uncached text request per document, up to 272,000 input tokens. Excludes storage, infrastructure, OCR, licensing, review, retries, tool calls, regional premiums, cache writes and processing-tier adjustments. Image tokenization is not modeled. Token counts change cost and, when TPM is set, quota-limited time. AI latency still uses the small-corpus benchmark. Azure quota examples keep direct OpenAI latency and prices; they are not Azure runtime or cost quotes. Test representative documents before budgeting.
Decisions handles configured choices, yes/no questions and scores. This planner's metadata example assigns a lifecycle stage from predefined values. Open-ended names, dates, summaries or written explanations need a text-generating endpoint such as Responses. The current task benchmark does not measure open-ended extraction.
Select Responses when a written classification note matters. Decisions does not return that written rationale. A separate explanation request would add time and cost and is not included here.
Prepared October 8, 2026. Published storage limits and API availability may change.
This calculator runs locally in your browser. It does not upload documents or call an AI service. Only the numeric settings and selected options are included in a copied settings link.