GPU and Server Configuration for Enterprise Local Drawing AI — 20-User Reference Cost
Published 2026-08-27 · Updated 2026-08-27 · Last Verified 2026-08-27
Answer
Enterprise local drawing AI requires a four-part consideration: GPU selection, server configuration, AI processing pipeline, and operations. GPU tiers range from RTX 5090 (dev PoC) to RTX PRO 6000 ×2 (20-user recommended) to B200 (large-scale / 100B+ models). The AI pipeline should process drawings through VLM → Drawing JSON/Geometry IR → CAD Engine → Validator, rather than expecting a VLM to generate precise CAD coordinates directly. GPU sizing is determined by VRAM capacity, model size, concurrent requests, and response time targets — not by a simple "users ÷ GPUs" formula. B200 is unnecessary for typical 20-user enterprise workloads and should only be considered for 100B+ models, multiple concurrent VLMs, fine-tuning, or hundreds of simultaneous drawing pages.
Applicability of This Document
| Product | Enterprise Local Drawing AI Server |
|---|---|
| Verified Versions | RTX 5090 / RTX PRO 6000 Blackwell / B200 / Mac M4 (as of August 2026) |
| Environment | On-premise (enterprise server room / colocation) |
| Required Permissions | None (read-only) |
| Execution Impact | None (read-only) |
| Restart | None |
| Last Verified | 2026-08-27 |
Ready-to-Run Commands
- 対象
- NVIDIA GPUs (RTX 5090 / RTX PRO 6000 / B100 / B200)
- 権限
- None (read-only)
- 変更作業
- None
- Production実行
- N/A
# Runnable model sizes by GPU VRAM (quantization: FP16 / INT4 approximate)
VRAM 24GB (RTX 5090 single)
FP16: 7B–13B (no quantization needed)
INT4: 14B–30B (loadable with quantization)
Use: Dev PoC, personal experiments, simple drawing recognition
VRAM 48GB (RTX 6000 Ada single)
FP16: 13B–30B (no quantization needed)
INT4: 30B–70B (loadable with quantization)
Use: Extended dev PoC, simple CAD drawing recognition
VRAM 96GB (RTX PRO 6000 Blackwell single)
FP16: 30B–50B (no quantization needed)
INT4: 50B–110B (loadable with quantization)
Use: Enterprise small-scale (1–5 users), Qwen3-VL-32B single
VRAM 192GB (B200 single)
FP16: 70B–100B (no quantization needed)
INT4: 100B–400B+ (loadable with quantization)
Use: Large-scale, multiple VLM concurrent, fine-tuning
Important: Model size is not half of VRAM requirement.
Example: Qwen3-VL-32B requires ~64GB at FP16 (32B parameters × 2bytes).
Quantization (INT4) reduces this to ~16GB, but with a precision/speed tradeoff.VRAM requirements scale with model size. Quantization reduces requirements but may degrade precision. Actual sizes vary by inference framework (llama.cpp / vLLM / Ollama).
- 対象
- GPU sizing (enterprise environment)
- 権限
- None (read-only)
- 変更作業
- None
- Production実行
- N/A
# GPU count calculation for 20 enterprise users (example)
【Conditions】
- Model: Qwen3-VL-32B (FP16: 64GB VRAM)
- Concurrent users: 20
- Max concurrent requests: 5
- VRAM per request: ~16GB (INT4 quantized)
- Response time target: under 30 seconds
- GPU VRAM per unit: 96GB (RTX PRO 6000 Blackwell)
【Calculation】
1) Concurrent requests per GPU
= 96GB / 16GB = 6 parallel
2) GPUs needed for 20 concurrent users
= ceil(5 / 6) = 1 ... theoretically 1 GPU suffices
3) In practice, add margin for:
- User growth
- Failover capacity
- Model update switching time
→ Recommended: 2 GPUs (replica A / replica B)
【Key point】
"20 users = 2 GPUs" is not universal.
Actual concurrent requests, drawing page count, resolution, and model size change the result.
PoC measurement is the only reliable method.GPU count is determined by concurrent request volume and VRAM per request — not by user count alone. Measure with PoC before finalizing.
How to Read the Results
| Column | Meaning | What to Check |
|---|---|---|
| Use Case | The use case this configuration suits | Corresponding GPU and memory configuration |
| GPU | Recommended GPU model | Memory capacity and compute performance |
| Memory | GPU memory or Unified Memory capacity | Runnable model size and quantization conditions |
| Model Reference | Model expected to run on this configuration | VRAM requirements and quantization conditions |
| Reference Use Case | Actual business application scenario | Concurrent users and throughput |
| Reference Cost | GPU unit price + full server (estimated as of August 2026) | Varies by exchange rate and supply conditions |
When to Use This
- Want to build local drawing AI but unsure of GPU selection criteria
- Cannot decide between RTX 5090 and RTX PRO 6000
- Want to know if Mac M4 or NVIDIA is better for drawing AI
- Do not know how many GPUs are needed for 20-user drawing AI
- Cannot determine if B200 is really necessary
- Want to know approximate cost for enterprise drawing AI setup
- Unclear on the role split between VLM and CAD Engine
Possible Causes (Most Likely First)
01
GPU selection is made "by price"
Choosing the most powerful GPU is not the right approach. Required VRAM depends on model size, quantization rate, and concurrent requests — meaning the most effective GPU varies by these factors.
02
The incorrect formula "users = GPU count"
Standard GPU sizing is based on VRAM capacity and model size, not user count. Even in a 20-user environment, one GPU may suffice depending on per-request workload.
03
VLM and CAD Engine role split is not understood
The VLM judges "what needs to change," while the CAD Engine generates "precise coordinates and shapes." Expecting the VLM to do both leads to coordinate precision failures.
04
B200 is being recommended beyond necessity
B200 is the most powerful GPU, but it is overkill for typical 20-user enterprise workloads. RTX PRO 6000 × 2 is sufficient in most cases and has better cost efficiency.
Diagnostic Steps
- 1
Analyze current drawing workload volume and frequency
Read-onlyAnalyze daily and monthly drawing processing volume, average pages per drawing, resolution, and format to estimate required VRAM and processing capacity.
- 2
Select target model size
Read-onlyReference: 4B–8B VLM for small-scale verification, 8B–32B for dev PoC, Qwen3-VL-32B or larger for production. Target model size determines required VRAM.
- 3
Measure actual VRAM consumption in a PoC environment
LowSend real concurrent requests with the selected model and measure VRAM usage, response time, and capacity. Measured values — not estimates — are the basis for final decisions.
Solutions
Low-risk action that can be done immediately
Measure actual user patterns in PoC
LowMeasure peak daily concurrent requests, drawing volume, and average processing time. Use these measured values to calculate GPU configuration. Measured data — not estimates — is the only reliable basis for sizing.
Separate the processing pipeline into VLM and CAD Engine
LowRather than "generating drawings with AI," structure it as "AI makes the judgment, CAD Engine generates the drawing." This simultaneously reduces VRAM requirements and improves precision.
Change requiring advance review
Expand GPU configuration in phases
MediumStart with RTX 5090 single GPU for PoC, then migrate to RTX PRO 6000 × 2 based on measured values. Do not purchase the final configuration upfront; expand based on PoC results.
Distribute requests via AI Gateway
MediumWith a 2-GPU configuration, use AI Gateway to distribute requests across replica A and B. This improves concurrent throughput and enables degraded operations during failures.
Establish monitoring infrastructure
MediumMonitor GPU utilization, VRAM usage, response time, and throughput. Use this for early detection of capacity constraints and as the basis for scaling decisions.
Work requiring expert review
Evaluate B200 adoption
Expert Review RequiredOnly consider B200 when ALL of the following apply: (1) 100B+ model is required, (2) multiple VLMs or agents simultaneously load the VRAM, (3) fine-tuning or large-scale training is run, (4) hundreds of drawing pages are processed across multiple projects simultaneously. RTX PRO 6000 × 2 is sufficient for typical 20-user enterprise workloads.
!Warnings
- GPU selection is determined by VRAM capacity, model size, concurrent requests, and response time targets — not by "users ÷ GPU count."
- "20 users = 2 GPUs" is an example, not a universal rule. Actual configuration varies based on PoC measurements.
- B200 is unnecessary for typical 20-user enterprise workloads. Only consider it for 100B+ models, multiple concurrent loads, or fine-tuning.
- Mac M4 Unified Memory has advantages for development and PoC, but NVIDIA is superior for production enterprise multi-user deployments due to CUDA absence.
- GPU prices fluctuate. Reference costs are estimates as of August 2026; confirm actual prices with suppliers.
- Expecting VLM to directly generate CAD drawings leads to coordinate precision failures. Clarify the role split.
Differences by Version and Environment
If This Doesn't Resolve It
Analyze actual drawing workload size distribution
Analyze average page count, resolution, and format per drawing to derive reference values for required VRAM capacity.
Measure peak concurrent request volume
Measure concurrent request peaks in the current (or similar) system to use as input for GPU configuration calculations.
Verify quantization precision impact
Confirm whether output precision with INT4 quantization meets business requirements in a PoC. FP16 may be needed for precision-critical use cases.
Basis and Limitations of This Document
Content verified through real-world operations
GPU comparison and cost estimates in this article are compiled from publicly available information as of August 2026. GPU prices and available models change; confirm actual costs with suppliers. VRAM requirements vary by inference framework.
FAQ
Should I choose RTX 5090 or RTX PRO 6000?
RTX 5090 (32GB) for dev PoC / personal experiments; RTX PRO 6000 Blackwell (96GB) for enterprise production. RTX 5090 prioritizes development speed, RTX PRO 6000 prioritizes VRAM capacity and stability.
Is 2 GPUs really necessary? Is 1 not enough?
The reason we recommend 2 GPUs is not that 1 cannot handle 20 users. It is for: failover capacity, user growth headroom, model update switching time, and future extensibility. A PoC measurement may show 1 GPU is sufficient, in which case start with 1 and expand as needed.
Is Mac M4 suitable for drawing AI?
Viable for development and PoC. Unified Memory capacity and MLX / llama.cpp / Ollama / Metal ecosystem are advantages. However, without CUDA support, NVIDIA is superior for production enterprise multi-user deployments. A strategy of starting with Mac and migrating to NVIDIA after PoC is also valid.
Is B200 really necessary?
Unnecessary for typical 20-user enterprise workloads. B200 is only justified for: 100B+ models, multiple concurrent VLMs, fine-tuning / large-scale training, or hundreds of simultaneous drawing pages. RTX PRO 6000 × 2 is sufficient outside these conditions.
What do GPUs cost?
Estimated as of August 2026: RTX 5090 (~USD 1,500–2,500), RTX PRO 6000 Blackwell single (~USD 6,000–8,000), RTX PRO 6000 × 2 (~USD 12,000–16,000), B200 (quotation required). GPU prices fluctuate; confirm with suppliers. Full server (CPU, RAM, storage, network) is additional.
How do VLM and CAD Engine split responsibilities?
The VLM handles judgment ("what parts need what changes"), the CAD Engine handles precise coordinate and shape generation. Expecting the VLM to generate precise CAD coordinates leads to precision failures. This separation simultaneously reduces VRAM requirements and improves precision.
Questions This Document Covers
- What is the configuration for enterprise drawing AI server?
- Which is better for drawing AI: RTX 5090 or RTX PRO 6000?
- How many GPUs needed for 20-user AI server?
- Can I build drawing AI with Mac M4?
- Is B200 really necessary?
What the Risk Labels Mean
- Read-onlyDoes not change data or settings.
- LowImpact is limited, but permissions and load should be checked.
- MediumMay affect performance, locking, or cost.
- HighMay cause an outage, data loss, or require recovery work.
- Expert Review RequiredRequires a separate review before applying to production.
GIIP's Scope of Support
The above is general sizing guidance not tied to specific GPU products. Regarding what GIIP covers: GIIP handles the full operations layer including GPU server design and selection, model deployment and optimization, AI Gateway-based request distribution, GPU utilization and model state monitoring, incident response, model version updates, cost optimization, and capacity expansion decisions. GPU procurement and installation are the customer's responsibility, but GIIP can support the surrounding software layer and operational design.
Author & Technical Review
GIIP Production Operations Team
About 30 years of experience in designing, migrating, and operating large-scale web services, SQL Server, Oracle, AWS, and Azure. Experience includes 12 sets of x12large-class AWS RDS for SQL Server environments, an Oracle environment with roughly 120,000 tables, and migrating roughly 3TB from TiDB to Aurora MySQL. AI agents and human experts currently continue to monitor and operate multiple cloud databases and about 30 web services.
How to Reduce Model Costs with an AI Router, and Failover Design for External AI API Outages
This article summarizes a design for reducing model costs through task-difficulty-based routing and quality gates, along with a failover design assuming multiple providers. It does not cover pricing or cost-reduction figures.
ai-operationsWhy AI automated execution needs approval and rollback, and how to design for it
Design elements for automated execution — snapshots, approval gates, dry-run separation, allowlists, idempotency, audit logs, staged rollout, and a kill switch — are summarized, along with where to draw the line on what can run without approval.
ai-operationsWhat it takes to bring an AI-generated application to production
We organize the remaining work between "the code works" and "it can run in production" into 14 items, with commands for detecting hardcoded keys and auditing dependency vulnerabilities.
giipWhat is the difference between a coding agent and GIIP FDE Ops?
Coding agents and operations services differ in their "unit of work." We compare the boundary between the two across nine dimensions, with commands you can run to check your own organization.
Related Services
GIIP can support GPU configuration validation through PoC with your actual drawings
If you need to repeat the same checks across multiple environments, we can discuss your entire operations setup.
GIIP can support GPU configuration validation through PoC with your actual drawings