giip
SES Proposal
GPU InfrastructureRead-onlyGPU図形AIVLMQwen3-VLローカルLLMCADRTX 5090RTX PRO 6000B200企業向けAIハードウェア構成コスト

GPU and Server Configuration for Enterprise Local Drawing AI — 20-User Reference Cost

Published 2026-08-27 · Updated 2026-08-27 · Last Verified 2026-08-27

Answer

Enterprise local drawing AI requires a four-part consideration: GPU selection, server configuration, AI processing pipeline, and operations. GPU tiers range from RTX 5090 (dev PoC) to RTX PRO 6000 ×2 (20-user recommended) to B200 (large-scale / 100B+ models). The AI pipeline should process drawings through VLM → Drawing JSON/Geometry IR → CAD Engine → Validator, rather than expecting a VLM to generate precise CAD coordinates directly. GPU sizing is determined by VRAM capacity, model size, concurrent requests, and response time targets — not by a simple "users ÷ GPUs" formula. B200 is unnecessary for typical 20-user enterprise workloads and should only be considered for 100B+ models, multiple concurrent VLMs, fine-tuning, or hundreds of simultaneous drawing pages.

Applicability of This Document

ProductEnterprise Local Drawing AI Server
Verified VersionsRTX 5090 / RTX PRO 6000 Blackwell / B200 / Mac M4 (as of August 2026)
EnvironmentOn-premise (enterprise server room / colocation)
Required PermissionsNone (read-only)
Execution ImpactNone (read-only)
RestartNone
Last Verified2026-08-27

Ready-to-Run Commands

Runnable Model Sizes by GPU VRAM (as of August 2026)Read-only
対象
NVIDIA GPUs (RTX 5090 / RTX PRO 6000 / B100 / B200)
権限
None (read-only)
変更作業
None
Production実行
N/A
# Runnable model sizes by GPU VRAM (quantization: FP16 / INT4 approximate)

VRAM 24GB (RTX 5090 single)
  FP16:   7B–13B (no quantization needed)
  INT4:   14B–30B (loadable with quantization)
  Use:    Dev PoC, personal experiments, simple drawing recognition

VRAM 48GB (RTX 6000 Ada single)
  FP16:   13B–30B (no quantization needed)
  INT4:   30B–70B (loadable with quantization)
  Use:    Extended dev PoC, simple CAD drawing recognition

VRAM 96GB (RTX PRO 6000 Blackwell single)
  FP16:   30B–50B (no quantization needed)
  INT4:   50B–110B (loadable with quantization)
  Use:    Enterprise small-scale (1–5 users), Qwen3-VL-32B single

VRAM 192GB (B200 single)
  FP16:   70B–100B (no quantization needed)
  INT4:   100B–400B+ (loadable with quantization)
  Use:    Large-scale, multiple VLM concurrent, fine-tuning

Important: Model size is not half of VRAM requirement.
Example: Qwen3-VL-32B requires ~64GB at FP16 (32B parameters × 2bytes).
Quantization (INT4) reduces this to ~16GB, but with a precision/speed tradeoff.

VRAM requirements scale with model size. Quantization reduces requirements but may degrade precision. Actual sizes vary by inference framework (llama.cpp / vLLM / Ollama).

How to Calculate GPU Count for 20 Enterprise UsersRead-only
対象
GPU sizing (enterprise environment)
権限
None (read-only)
変更作業
None
Production実行
N/A
# GPU count calculation for 20 enterprise users (example)

【Conditions】
- Model: Qwen3-VL-32B (FP16: 64GB VRAM)
- Concurrent users: 20
- Max concurrent requests: 5
- VRAM per request: ~16GB (INT4 quantized)
- Response time target: under 30 seconds
- GPU VRAM per unit: 96GB (RTX PRO 6000 Blackwell)

【Calculation】
1) Concurrent requests per GPU
   = 96GB / 16GB = 6 parallel

2) GPUs needed for 20 concurrent users
   = ceil(5 / 6) = 1 ... theoretically 1 GPU suffices

3) In practice, add margin for:
   - User growth
   - Failover capacity
   - Model update switching time
   → Recommended: 2 GPUs (replica A / replica B)

【Key point】
"20 users = 2 GPUs" is not universal.
Actual concurrent requests, drawing page count, resolution, and model size change the result.
PoC measurement is the only reliable method.

GPU count is determined by concurrent request volume and VRAM per request — not by user count alone. Measure with PoC before finalizing.

How to Read the Results

ColumnMeaningWhat to Check
Use CaseThe use case this configuration suitsCorresponding GPU and memory configuration
GPURecommended GPU modelMemory capacity and compute performance
MemoryGPU memory or Unified Memory capacityRunnable model size and quantization conditions
Model ReferenceModel expected to run on this configurationVRAM requirements and quantization conditions
Reference Use CaseActual business application scenarioConcurrent users and throughput
Reference CostGPU unit price + full server (estimated as of August 2026)Varies by exchange rate and supply conditions

When to Use This

  • Want to build local drawing AI but unsure of GPU selection criteria
  • Cannot decide between RTX 5090 and RTX PRO 6000
  • Want to know if Mac M4 or NVIDIA is better for drawing AI
  • Do not know how many GPUs are needed for 20-user drawing AI
  • Cannot determine if B200 is really necessary
  • Want to know approximate cost for enterprise drawing AI setup
  • Unclear on the role split between VLM and CAD Engine

Possible Causes (Most Likely First)

  1. 01

    GPU selection is made "by price"

    Choosing the most powerful GPU is not the right approach. Required VRAM depends on model size, quantization rate, and concurrent requests — meaning the most effective GPU varies by these factors.

  2. 02

    The incorrect formula "users = GPU count"

    Standard GPU sizing is based on VRAM capacity and model size, not user count. Even in a 20-user environment, one GPU may suffice depending on per-request workload.

  3. 03

    VLM and CAD Engine role split is not understood

    The VLM judges "what needs to change," while the CAD Engine generates "precise coordinates and shapes." Expecting the VLM to do both leads to coordinate precision failures.

  4. 04

    B200 is being recommended beyond necessity

    B200 is the most powerful GPU, but it is overkill for typical 20-user enterprise workloads. RTX PRO 6000 × 2 is sufficient in most cases and has better cost efficiency.

Diagnostic Steps

  1. 1

    Analyze current drawing workload volume and frequency

    Read-only

    Analyze daily and monthly drawing processing volume, average pages per drawing, resolution, and format to estimate required VRAM and processing capacity.

  2. 2

    Select target model size

    Read-only

    Reference: 4B–8B VLM for small-scale verification, 8B–32B for dev PoC, Qwen3-VL-32B or larger for production. Target model size determines required VRAM.

  3. 3

    Measure actual VRAM consumption in a PoC environment

    Low

    Send real concurrent requests with the selected model and measure VRAM usage, response time, and capacity. Measured values — not estimates — are the basis for final decisions.

Solutions

Low-risk action that can be done immediately

  • Measure actual user patterns in PoC

    Low

    Measure peak daily concurrent requests, drawing volume, and average processing time. Use these measured values to calculate GPU configuration. Measured data — not estimates — is the only reliable basis for sizing.

  • Separate the processing pipeline into VLM and CAD Engine

    Low

    Rather than "generating drawings with AI," structure it as "AI makes the judgment, CAD Engine generates the drawing." This simultaneously reduces VRAM requirements and improves precision.

Change requiring advance review

  • Expand GPU configuration in phases

    Medium

    Start with RTX 5090 single GPU for PoC, then migrate to RTX PRO 6000 × 2 based on measured values. Do not purchase the final configuration upfront; expand based on PoC results.

  • Distribute requests via AI Gateway

    Medium

    With a 2-GPU configuration, use AI Gateway to distribute requests across replica A and B. This improves concurrent throughput and enables degraded operations during failures.

  • Establish monitoring infrastructure

    Medium

    Monitor GPU utilization, VRAM usage, response time, and throughput. Use this for early detection of capacity constraints and as the basis for scaling decisions.

Work requiring expert review

  • Evaluate B200 adoption

    Expert Review Required

    Only consider B200 when ALL of the following apply: (1) 100B+ model is required, (2) multiple VLMs or agents simultaneously load the VRAM, (3) fine-tuning or large-scale training is run, (4) hundreds of drawing pages are processed across multiple projects simultaneously. RTX PRO 6000 × 2 is sufficient for typical 20-user enterprise workloads.

!Warnings

  • GPU selection is determined by VRAM capacity, model size, concurrent requests, and response time targets — not by "users ÷ GPU count."
  • "20 users = 2 GPUs" is an example, not a universal rule. Actual configuration varies based on PoC measurements.
  • B200 is unnecessary for typical 20-user enterprise workloads. Only consider it for 100B+ models, multiple concurrent loads, or fine-tuning.
  • Mac M4 Unified Memory has advantages for development and PoC, but NVIDIA is superior for production enterprise multi-user deployments due to CUDA absence.
  • GPU prices fluctuate. Reference costs are estimates as of August 2026; confirm actual prices with suppliers.
  • Expecting VLM to directly generate CAD drawings leads to coordinate precision failures. Clarify the role split.

Differences by Version and Environment

GPU model selection guide (as of August 2026)RTX 5090: dev PoC / personal experiments. RTX PRO 6000 Blackwell: enterprise single-GPU standard. RTX PRO 6000 × 2: 20-user recommended. B200: large-scale / special use cases.
Model size trendsQwen3-VL-32B requires ~64GB VRAM at FP16. Quantization (INT4) reduces this to ~16GB with a precision/speed tradeoff. Model size and VRAM requirements are proportional.
Mac M4 positioningMac M4 Max (128GB Unified Memory) is a viable option for development and PoC. MLX / llama.cpp / Ollama / Metal ecosystem is available. However, without CUDA support, NVIDIA is superior for production enterprise multi-user deployments.

If This Doesn't Resolve It

  • Analyze actual drawing workload size distribution

    Analyze average page count, resolution, and format per drawing to derive reference values for required VRAM capacity.

  • Measure peak concurrent request volume

    Measure concurrent request peaks in the current (or similar) system to use as input for GPU configuration calculations.

  • Verify quantization precision impact

    Confirm whether output precision with INT4 quantization meets business requirements in a PoC. FP16 may be needed for precision-critical use cases.

Basis and Limitations of This Document

Content verified through real-world operations

GPU comparison and cost estimates in this article are compiled from publicly available information as of August 2026. GPU prices and available models change; confirm actual costs with suppliers. VRAM requirements vary by inference framework.

FAQ

Should I choose RTX 5090 or RTX PRO 6000?

RTX 5090 (32GB) for dev PoC / personal experiments; RTX PRO 6000 Blackwell (96GB) for enterprise production. RTX 5090 prioritizes development speed, RTX PRO 6000 prioritizes VRAM capacity and stability.

Is 2 GPUs really necessary? Is 1 not enough?

The reason we recommend 2 GPUs is not that 1 cannot handle 20 users. It is for: failover capacity, user growth headroom, model update switching time, and future extensibility. A PoC measurement may show 1 GPU is sufficient, in which case start with 1 and expand as needed.

Is Mac M4 suitable for drawing AI?

Viable for development and PoC. Unified Memory capacity and MLX / llama.cpp / Ollama / Metal ecosystem are advantages. However, without CUDA support, NVIDIA is superior for production enterprise multi-user deployments. A strategy of starting with Mac and migrating to NVIDIA after PoC is also valid.

Is B200 really necessary?

Unnecessary for typical 20-user enterprise workloads. B200 is only justified for: 100B+ models, multiple concurrent VLMs, fine-tuning / large-scale training, or hundreds of simultaneous drawing pages. RTX PRO 6000 × 2 is sufficient outside these conditions.

What do GPUs cost?

Estimated as of August 2026: RTX 5090 (~USD 1,500–2,500), RTX PRO 6000 Blackwell single (~USD 6,000–8,000), RTX PRO 6000 × 2 (~USD 12,000–16,000), B200 (quotation required). GPU prices fluctuate; confirm with suppliers. Full server (CPU, RAM, storage, network) is additional.

How do VLM and CAD Engine split responsibilities?

The VLM handles judgment ("what parts need what changes"), the CAD Engine handles precise coordinate and shape generation. Expecting the VLM to generate precise CAD coordinates leads to precision failures. This separation simultaneously reduces VRAM requirements and improves precision.

Questions This Document Covers

  • What is the configuration for enterprise drawing AI server?
  • Which is better for drawing AI: RTX 5090 or RTX PRO 6000?
  • How many GPUs needed for 20-user AI server?
  • Can I build drawing AI with Mac M4?
  • Is B200 really necessary?

What the Risk Labels Mean

  • Read-onlyDoes not change data or settings.
  • LowImpact is limited, but permissions and load should be checked.
  • MediumMay affect performance, locking, or cost.
  • HighMay cause an outage, data loss, or require recovery work.
  • Expert Review RequiredRequires a separate review before applying to production.

GIIP's Scope of Support

The above is general sizing guidance not tied to specific GPU products. Regarding what GIIP covers: GIIP handles the full operations layer including GPU server design and selection, model deployment and optimization, AI Gateway-based request distribution, GPU utilization and model state monitoring, incident response, model version updates, cost optimization, and capacity expansion decisions. GPU procurement and installation are the customer's responsibility, but GIIP can support the surrounding software layer and operational design.

Author & Technical Review

GIIP Production Operations Team

About 30 years of experience in designing, migrating, and operating large-scale web services, SQL Server, Oracle, AWS, and Azure. Experience includes 12 sets of x12large-class AWS RDS for SQL Server environments, an Oracle environment with roughly 120,000 tables, and migrating roughly 3TB from TiDB to Aurora MySQL. AI agents and human experts currently continue to monitor and operate multiple cloud databases and about 30 web services.

Related Knowledge

Related Services

GIIP can support GPU configuration validation through PoC with your actual drawings

If you need to repeat the same checks across multiple environments, we can discuss your entire operations setup.

GIIP can support GPU configuration validation through PoC with your actual drawings

← Back to Knowledge Base