100 Gbit Windows Servers are live — launch-sale discounts for early adoptersExplore products →
High-performance server infrastructure

EPYC servers on 100 Gbit.

Dedicated AMD EPYC cores, PCIe NVMe storage, 100 Gbit networking, and full administrator access. No shared tenancy.

Dedicated resourcesLatest-generation computeWindows or Ubuntu
Products

Four products. One dedicated stack.

Choose a ready-to-run Agent OS, dedicated Windows compute, unified model access, or a private GPU lane.

Agent OS servers

Six Agent OS tiers. Pick your scale.

Dedicated compute, unmetered local inference, and cloud model access on one server, sized from a first production deployment to multi-fleet orchestration.

Monster

Scale
$350/ month

More RAM to run more agents per box.

Server specs

96 GBRAM
32Cores
2 TBNVMe

Everything in EXTREME, and also:

  • Up to 500 TPS local inference
  • 500M pooled cloud tokens
  • Five additional cloud models
18 MODELS INCLUDEDUnlocks Luna, DeepSeek, Qwen, Devstral, and gpt-oss.
Unlimited local · 100 Gbit
Configure Monster

Beast

Production
$450/ month

For production fleets running agents in parallel.

Server specs

128 GBRAM
40Cores
2 TBNVMe

Everything in MONSTER, and also:

  • Up to 1,000 TPS local inference
  • 1B pooled cloud tokens
  • Three premium generalists
21 MODELS INCLUDEDAdds GPT-6 Sol and both MiniMax tiers.
Unlimited local · 100 Gbit
Configure Beast

Lightning

Speed
$500/ month

High RAM for many concurrent agent sessions.

Server specs

256 GBRAM
48Cores
2 TBNVMe

Everything in BEAST, and also:

  • Up to 2,000 TPS local inference
  • 1.25B pooled cloud tokens
  • Both Grok 4.1 Fast modes
23 MODELS INCLUDEDFast reasoning and non-reasoning Grok lanes.
Unlimited local · 100 Gbit
Configure Lightning

Dragon

Frontier
$550/ month

Memory headroom for long-context, multi-step workflows.

Server specs

512 GBRAM
48Cores
4 TBNVMe

Everything in LIGHTNING, and also:

  • Up to 4,000 TPS local inference
  • 1.5B pooled cloud tokens
  • Two premier Grok models
25 MODELS INCLUDEDAdds Grok 4.3 and Grok 4.20 non-reasoning.
Unlimited local · 100 Gbit
Configure Dragon

Megatron

Maximum
$600/ month

All 26 pooled models and the largest local footprint.

Server specs

1 TBRAM
48Cores
4 TBNVMe

Everything in DRAGON, and also:

  • Up to 10,000 TPS local inference
  • 2B pooled cloud tokens
  • $250 Grok 4.6 flagship credit
ALL 26 POOLED MODELSAdds Grok 4.20 reasoning and flagship credit.
Unlimited local · 100 Gbit
Configure Megatron
Model access by plan

All plans include

Every Agent OS server starts with these 13 pooled models. Higher tiers keep every model below and add the models listed in their tier.

GPT-5 nanoGPT-4.1 nanoGPT-4o Minigpt-oss-20bQwen 3.6 27BQwen 3.8 27BQwen3 Coder NextGemma 4 31BDeepSeek V4 FlashDeepSeek V4 Flash 0731GLM 5.3 FlashAgents-A1Nemotron 3.5 30B A3B

Additional models unlocked

Monster · +5

GPT-6 Luna
DeepSeek V4.1 Flash
Qwen 3.6 35B-A3B
Devstral 2 123B
gpt-oss-120b

Beast · +3

GPT-6 Sol
MiniMax M2.7
MiniMax M3

Lightning · +2

Grok 4.1 Fast · reasoning
Grok 4.1 Fast · non-reasoning

Dragon · +2

Grok 4.3
Grok 4.20 · non-reasoning

Megatron · +1

Grok 4.20 · reasoning

All 26 pooled models

Flagship credit is separate from the pooled base-model allowance.

Windows servers

Windows servers on AMD EPYC.

Dedicated AMD EPYC cores, PCIe NVMe, administrator access, and a 100 Gbit network on every plan.

Starter

Windows
$39.99/ month

Entry server for a few bots or scheduled scripts.

8 GBRAM
4Cores
100 GBNVMe
  • 4.1 GHz CPU cores
  • Administrator access
  • 100 Gbit network
Dedicated resources
Order Starter

Premium

Windows
$49.99/ month

More cores and RAM for always-on daily jobs.

16 GBRAM
8Cores
100 GBNVMe
  • 4.1 GHz CPU cores
  • Administrator access
  • 100 Gbit network
Dedicated resources
Order Premium

Ultimate

Windows
$69.99/ month

Even CPU-to-RAM split for steady multi-bot loads.

24 GBRAM
12Cores
100 GBNVMe
  • 4.1 GHz CPU cores
  • Administrator access
  • 100 Gbit network
Dedicated resources
Order Ultimate

VIP

Windows
$89.99/ month

Extra cores for running tasks concurrently.

32 GBRAM
18Cores
100 GBNVMe
  • 4.1 GHz CPU cores
  • Administrator access
  • 100 Gbit network
Dedicated resources
Order VIP

Extreme

Windows
$104.99/ month

High RAM for browser farms and local models.

48 GBRAM
24Cores
100 GBNVMe
  • 4.1 GHz CPU cores
  • Administrator access
  • 100 Gbit network
Dedicated resources
Order Extreme

Monster

Windows
$149.99/ month

Top CPU tier with double the NVMe storage.

64 GBRAM
32Cores
200 GBNVMe
  • 4.1 GHz CPU cores
  • Administrator access
  • 100 Gbit network
Dedicated resources
Order Monster
API plans + VPS

One endpoint. Your whole model stack.

A scoped LiteLLM key paired with dedicated VPS resources, so compute and model access arrive together.

Monster API

18 models
$149/ month

Adds models and VPS resources for more agents.

16 GBRAM
8Cores
100 GBNVMe

Everything in EXTREME API, and also:

  • 500M pooled cloud tokens
  • Five additional models
  • More concurrent headroom
VPS + API · 100 Gbit
Ask about Monster API

Beast API

21 models
$249/ month

Larger model pool and VPS for production.

24 GBRAM
12Cores
100 GBNVMe

Everything in MONSTER API, and also:

  • 1B pooled cloud tokens
  • GPT-6 Sol
  • MiniMax M2.7 and M3
VPS + API · 100 Gbit
Ask about Beast API

Lightning API

23 models
$299/ month

Low-latency models for high-concurrency agent runs.

32 GBRAM
18Cores
100 GBNVMe

Everything in BEAST API, and also:

  • 1.25B pooled cloud tokens
  • Both Grok 4.1 Fast modes
  • Expanded concurrency
VPS + API · 100 Gbit
Ask about Lightning API

Dragon API

25 models
$399/ month

Wider model catalog and more RAM for larger fleets.

48 GBRAM
24Cores
100 GBNVMe

Everything in LIGHTNING API, and also:

  • 1.5B pooled cloud tokens
  • Grok 4.3
  • Grok 4.20 non-reasoning
VPS + API · 100 Gbit
Ask about Dragon API

Megatron API

26 pooled
$499/ month

All 26 pooled models, plus a dedicated VPS.

64 GBRAM
32Cores
200 GBNVMe

Everything in DRAGON API, and also:

  • 2B pooled cloud tokens
  • Grok 4.20 reasoning
  • $250 Grok 4.6 credit
VPS + API · 100 Gbit
Ask about Megatron API

Pooled model allowances reset monthly. Flagship models bill separately against metered credit.

Dedicated GPU rental

Dedicated GPUs. Not time-sliced.

Reserve a single GPU today. For multi-GPU deployments, scope a custom build with the team.

Dual GPU

Custom
Quote/ scope

Two GPUs in one dedicated node for larger models.

  • Higher aggregate throughput
  • Workload placement review
  • Full root access
  • Subject to inventory
Dual GPU · custom configuration
Request a dual-GPU quote

GPU Cluster

Custom
Quote/ scope

Reserved GPU capacity spread across multiple nodes.

  • Private network fabric
  • Deployment planning
  • Capacity sized to workload
  • Custom commitment term
Multi-node · reserved capacity
Plan a GPU cluster
Models

Every model, mapped to its tier.

Browse 26 pooled models by the first Agent OS tier that includes them, plus four flagship models billed through metered credit.

30 models

Pooled models draw from your monthly base allowance. Grok 4.6, qwen3.8-max, kimi-k3, and GPT-6 Astra are metered flagship lanes, billed against separate credit.

How it works

Local for volume. Cloud for specialists.

Routine work stays on the unmetered local lane. Specialist calls route through the API catalog included with your tier.

01

Choose a product

Pick Agent OS, Windows, API + VPS, or dedicated GPU capacity.

02

Get your environment

Server resources, credentials, and API access are provisioned to match your tier.

03

Run local volume

Run recurring agent loops on your dedicated local inference lane. No per-token billing.

04

Route specialist calls

Call cloud models with your OpenAI-compatible key when a task needs them.

Plan Benefits

Tuned at every layer, from BIOS to runtime.

Your dedicated RTX PRO 6000 environment comes pre-installed with Hermes and OpenClaw out of the box. Enjoy dedicated privacy, live state snapshots, and zero rate limits across your optimized Agent OS.

01

Preloaded Agent Stack

Windows or Ubuntu, preloaded with OpenWebUI, Hermes, and OpenClaw. Log in and deploy your swarms.

02

Unmetered Local Tokens

Burst past 10,000 TPS on RTX PRO 6000 hardware. No rate limits, no API queues, no per-token fees.

03

Live Snapshots

Agent state is fragile. Automatic environment snapshots let you roll back to a known-good state when a swarm crashes.

04

Bare-Metal Isolation

Your weights, your prompts, your data. Bare-metal isolation means no shared tenants and no third-party access to your workloads.

05

Long-Context Memory

ECC memory paired with GPU acceleration, sized to hold long context windows in memory.

06

100 Gbps Backbone

Direct uplinks to 100 Gbps transit. Your agents can crawl, pull, and ingest data at scale without saturating the pipe.

07

PCIe NVMe

Up to 14,000 MB/s reads and writes. Model weights load fast, and storage stops being the bottleneck.

08

Zen 4 Architecture

AMD EPYC™ 9274F processors with up to 48 cores, 4.3 GHz boost clocks, and large L3 cache for parallel agent workloads.

09

Locked-In Pricing

The price, speeds, and limits on your plan today stay fixed for life. No forced migrations. No surprise changes.

10

100% Uptime SLA

Covered by a 100% uptime SLA from our enterprise datacenter. Your fleet stays online 24/7.

11

BIOS & OS Tuned

Internal OS optimization scripts and custom BIOS settings, tuned for system, GPU, and network throughput.

12

Rolling Model Upgrades

When new frontier models ship, we add them to your cluster. No migration and no plan change required.

Questions, answered

Questions to settle before you order.

Setup, model access, infrastructure, and support, answered before you pick a plan.

What is included with an Agent OS server?

Each plan combines dedicated server resources, a ready-to-run Agent OS environment, unmetered local inference, and an OpenAI-compatible API key for the cloud models included with that tier.

Can I choose Windows or Ubuntu?

Yes. Agent OS environments can be delivered on Windows or Ubuntu, with administrator-level access to your dedicated resources.

How does model access change between tiers?

Every Agent OS plan includes the 13-model core catalogue. Each higher tier keeps every model below it and adds coding, reasoning, speed, and frontier lanes. The Models section lists the entry tier for each model.

Is the API OpenAI-compatible?

Yes. The included API is OpenAI-compatible. Point your existing SDKs, tools, and agent frameworks at it with a new base URL and key.

What does pooled token allowance mean?

Pooled cloud models draw from the monthly allowance included with your plan. Metered flagship models use explicit credit separately, so premium usage stays visible and controlled.

Can you build a custom GPU or multi-server configuration?

Yes. Dedicated GPUs, multi-GPU systems, blade deployments, and larger reserved environments can be scoped directly with the infrastructure team. Start a custom quote.

Who handles support?

EbotServers handles support directly: provisioning, infrastructure, networking, and account questions.

Need something more custom?

Spec a custom build for your fleet.

Legacy machine, multi-GPU cluster, blade system, or full cabinet. Talk directly with the team about capacity, networking, and deployment.

Contact us