Skip to main content
STATUS: PRE-CONSTRUCTION · SITE A UNDER EXCLUSIVITYNODE: VASILIKOS-01 — 34.7246°N, 33.2247°ECAMPUS: RISC-V PHASE 1 · MULTI-SILICON EVAL · PLANNEDPOWER: 42MW ON-SITE GENERATION · DESIGN TARGETSTATUS: PRE-CONSTRUCTION · SITE A UNDER EXCLUSIVITYNODE: VASILIKOS-01 — 34.7246°N, 33.2247°ECAMPUS: RISC-V PHASE 1 · MULTI-SILICON EVAL · PLANNEDPOWER: 42MW ON-SITE GENERATION · DESIGN TARGET
AGICY.AI
StackTechnology OverviewRISC-V sovereign stackComputeBare-metal EU compute
FacilitiesData CentersWorldwide map & trackerVasilikos Campus42MW sovereign campus briefingSustainability100% renewable mission
HardwareHardware FleetVendor hub · phased COD roadmapTenstorrent GalaxyPhase 1 Blackhole fleet (pre-COD)AESOLAR AlpineEnergy stack · hail-class PV + BESSAMD HeliosPhase 2 open rack-scale (eval · pre-COD)CerebrasPhase 2 · wafer-scale eval (pre-COD)
Submit your Hardware for reviewPropose accelerators for the Vasilikos fleet
AIAI Web SearchNew!Sovereign AI search · AGICY & the webAI My MapsNew!Smart Maps · voice routing · agentsAgents Trading CryptoBeta · New!Live paper-trading arena · BTC · ETH · SUIAI Video SearchNew!Find AI videos · avatars · films · adsAdvertiseSearch ad program · how it works · board ranks
GatewayCopperwayEU-sovereign OpenAI-compatible gatewayTry PlaygroundNewLive Copperway demo · PII vaultSovereign Exchange5-year cross-org sovereign plan
ProductsReserve CapacityPre-construction LOI tiersMarketplaceCompute marketplaceGPUs Rent LiveLiveEU partner GPU now · until CODCompute VouchersSovereign compute creditsModel LeaderboardFrontier model rankingsPricingSubscription tiers
WorkloadsComputeBare-metal EU inference & trainingCopperway GatewayOpenAI-compatible EU APIVasilikos Campus42MW sovereign campus briefing
Trust & complianceTrust CenterSecurity portal · docs · statusEU AI ActRegulatory mapping & controlsAI Readiness AuditPublic-data readiness hub
PricingTiers
Capital & EducationInvestInstitutional data room & deal flowAcademyAI training programs
Individuals & Family OfficesLiving in EUNewClass B capital allocation · no visa framingInternationalNewPlan B · equity alternative to property
IntelligenceResearchPublications & portals
CompanyAboutBrand · HoldCo targetMissionCharter & sovereigntyTrust CenterSecurity portal · docs · status
Schedule Briefing
Sign In
§ 00ServerlessOverview§ 01BenefitsWhy serverless§ 02PricingPay per token§ 03Use casesIdeal workloads§ 04FAQSupport§ 05StartDeploy

§ 01 Serverless · Overview

Serverless AI
Pay-Per-Token, Scale-to-Zero

Run sovereign AI inference without managing infrastructure. AGICY's serverless platform automatically scales your inference endpoints from zero to thousands of requests per second, charging you only for the tokens you actually process. Sub-100ms cold starts, no idle compute costs, and full GDPR compliance on EU-sovereign RISC-V infrastructure.

Scale-to-Zero
Zero
Pay-Per-Token
Per-Tok
Cold Start
<100ms
SLA
99.9%

§ 02 Benefits · Why serverless

Why Serverless AI

Eliminate infrastructure overhead, reduce costs, and ship AI-powered features faster — all on sovereign infrastructure.

Cost Efficiency

Pay only for the tokens you process — not for idle GPU hours. AGICY's serverless platform eliminates the cost of provisioned compute that sits unused between requests. For bursty workloads, this translates to Significant cost savings compared to dedicated instances. Your endpoints scale to zero during quiet periods, which means zero compute costs overnight, on weekends, or during seasonal lulls. When traffic spikes, the platform scales automatically without any manual intervention or pre-warming. There are no minimum commitments, no reserved capacity fees, and no bandwidth charges for inference traffic within AGICY's sovereign network. Usage-based billing with transparent per-token pricing gives you complete cost predictability and control.

60–80% Savings

Infinite Scaling

AGICY's serverless infrastructure handles everything from a single request per day to sustained bursts of thousands of concurrent inference requests — without any configuration changes on your part. The platform's intelligent request router distributes inference workloads across available Tenstorrent Galaxy servers in real-time, spinning up additional compute capacity within milliseconds when demand increases. Sub-100ms cold starts are achieved through our pre-warmed model pool architecture, where frequently-used models are kept in a ready state across the fleet. For less common models, warm-up happens transparently during the first request with no perceptible delay. Auto-scaling is fully sovereign — all scaling decisions are made locally on EU infrastructure.

Auto-Scale

Developer Simplicity

No servers to provision. No clusters to manage. No capacity to plan. AGICY's serverless platform abstracts away the entire infrastructure layer so your engineering team can focus exclusively on building AI-powered features. Deploy a new inference endpoint with a single API call, switch between models instantly, and iterate at the speed of your product roadmap — not your infrastructure team's sprint cycle. Our OpenAI-compatible API means you can migrate existing applications to sovereign serverless with a single base URL change. Full support for streaming, function calling, tool use, and JSON mode — all handled automatically by the serverless runtime. Comprehensive observability with per-request metrics, cost tracking, and latency percentiles is included out of the box.

Zero-Ops

§ 03 Pricing · Pay per token

Serverless Pricing Model

Transparent per-token pricing with no hidden fees. Pay only for what you use, scale to zero when idle.

Model TierInput TokensOutput Tokens
Small (7B–13B params)€0.10 / 1M tokens€0.30 / 1M tokens
Medium (30B–70B params)€0.50 / 1M tokens€1.50 / 1M tokens
Large (100B+ params)€2.00 / 1M tokens€6.00 / 1M tokens
Embedding Models€0.05 / 1M tokens—
Idle Compute€0 (scale-to-zero)€0 (scale-to-zero)
Network Egress€0 (scale-to-zero)€0 (scale-to-zero)
Minimum CommitmentNoneNone

All prices are in euros and include EU data sovereignty, GDPR compliance, and full audit trail transparency. Volume discounts are available for organisations processing more than 100M tokens per month. Enterprise plans with committed-use discounts and custom SLAs are available through our pricing page. Startups and academic institutions may qualify for credits through our Startup Program.

§ 04 Use Cases · Ideal workloads

Ideal Use Cases

Serverless AI is the perfect fit for workloads with variable demand, tight budgets, or rapid prototyping requirements.

Startups & MVPs

Launch AI-powered products without upfront infrastructure investment. Serverless lets you validate product-market fit with real AI capabilities at minimal cost. Start with a handful of requests per day during beta, then scale seamlessly to thousands of concurrent users as your product grows. No need to renegotiate infrastructure contracts or provision additional capacity — the platform handles growth automatically. Combined with AGICY's Startup Program, you can access sovereign AI inference with generous free-tier credits.

Low Entry Cost
🔄

Event-Driven Workloads

Process AI workloads triggered by events — user actions, webhook callbacks, scheduled jobs, or real-time data streams. Serverless endpoints activate instantly when events arrive and return to zero when the event stream pauses. This makes serverless ideal for chatbots that experience peak traffic during business hours, content moderation pipelines that run on publish events, document processing triggered by uploads, and notification systems that generate personalised AI summaries. Pay-per-token pricing ensures you only incur costs when actual work is being performed.

Event-Triggered
🧪

Experimentation & A/B Testing

Compare multiple models, prompts, and configurations without provisioning separate infrastructure for each variant. AGICY's serverless platform lets you deploy dozens of model variants simultaneously — each behind its own endpoint — and route traffic between them for A/B testing and experimentation. Test Llama 4 Maverick against Mistral Large 2 on your production queries, compare different prompt engineering strategies, or evaluate fine-tuned models against their base versions. Each endpoint scales independently and is billed only for its actual token usage, making experimentation virtually free during low-traffic periods.

Multi-Model

§ 05 Questions · FAQ

Frequently Asked Questions

Common questions about serverless AI inference on AGICY sovereign infrastructure.

How fast are cold starts on the serverless platform?
AGICY's serverless platform achieves sub-100ms cold starts for the most popular models through our pre-warmed model pool architecture. Frequently requested models like Llama 4 Maverick, Mistral Large 2, and DeepSeek V3 are maintained in a ready state across the Tenstorrent Galaxy fleet, meaning the first request after an idle period experiences no perceptible cold start penalty. For less common or custom fine-tuned models, cold starts typically complete within 200–500ms as the model weights are loaded from NVMe storage into RISC-V memory. Subsequent requests to a warm endpoint experience the same sub-50ms time-to-first-token latency as dedicated inference endpoints.
How does pay-per-token pricing work exactly?
You are billed for every token processed by the inference engine — both input tokens (your prompt and context) and output tokens (the model's response). Pricing varies by model size tier, with small models (7B–13B parameters) starting at €0.10 per million input tokens and Contact for pricing output tokens. There are no charges for idle time, network egress, or API calls that don't produce tokens. Billing is calculated per-request and aggregated into your monthly invoice with detailed breakdowns by model, endpoint, and time period. Real-time usage dashboards and budget alerts help you monitor and control spend. Volume discounts apply automatically above 100M tokens per month. There are no minimum commitments or reserved capacity fees.
Are there any rate limits or scaling constraints?
Default rate limits are set at 1,000 requests per minute and 100,000 tokens per minute per endpoint, which is sufficient for most production workloads. These limits are soft limits and can be increased upon request — enterprise customers routinely operate at 10,000+ requests per minute with custom rate limit configurations. The platform's auto-scaling architecture supports bursts up to 50× the baseline request rate with no degradation in latency, as new compute capacity is provisioned within milliseconds from the pre-warmed pool. There are no hard upper bounds on total throughput; the platform scales horizontally across the entire Tenstorrent Galaxy fleet as needed. Contact our team to discuss custom scaling configurations.
Is serverless AI still sovereign?
Yes — AGICY's serverless platform provides the same data sovereignty guarantees as our dedicated and on-premises deployment modes. Every serverless inference request is processed on RISC-V silicon physically located in AGICY's EU-sovereign facility in Cyprus. Your prompts and completions never leave EU jurisdiction. The serverless control plane — including the request router, auto-scaler, and billing system — is also hosted entirely on EU-sovereign infrastructure with no dependencies on US cloud providers. There is zero exposure to the US CLOUD Act because AGICY is a Cyprus (EU) entity. We don't log your prompts, don't train on your data, and provide full audit trail transparency through our See our Trust Center for full audit trail documentation.

§ 06 Get Started · Deploy

Go Serverless

Start deploying AI models with pay-per-token pricing and scale-to-zero efficiency. No infrastructure. No idle costs. Full sovereignty.

Go Serverless →View Pricing
solutions/inferencepricingstartup-programcompute

§ FIN — Close of Document

Ready to build on sovereign infrastructure?

Schedule a confidential briefing with our team. NDA-protected, no commitment.

Schedule a briefing →
EU JURISDICTION · CYPRUSGDPR ART. 28 DPA-READY · BY DESIGNNIS2-ALIGNED · BY DESIGNEU AI ACT ART. 12 LOGGING SUPPORT · BY DESIGNRISC-V NATIVE · OPEN ISA

Design-alignment statements for a pre-construction facility — not certifications or attestations. Basis: compliance FAQ, § 08. Careers: we aim for 50-50 gender balance across hiring cohorts.

AGICY.AI

Advanced Governance & Intelligence Cyprus

The sovereign architecture for the Cyprus mind.
Humanitarian mandate: civilian public benefit only — healthcare, education, civil resilience. Civilian / humanitarian mandate only.
Office: 8 John Kennedy Street, Iris House, 7th floor, 3106 Limassol, Cyprus
+357 95 572 777 · 08:30 – 19:00 · agi@agicy.ai
VASILIKOS ENERGY CENTRE, LIMASSOL DISTRICT · PRE-CONSTRUCTION
34.7246°N · 33.2247°E
Principal campus: Cyprus Vasilikos (Phase 1). Parallel HoldCo path: sovereign compute project in Greece (TARGET / planning) — ~20 MW-class Tenstorrent / air-cooled inference positioning for EU diversification; separate CapEx, no offtake claimed.

The Ledger — monthly briefing
  • CopperwayEU-sovereign OpenAI-compatible gateway
  • Try PlaygroundNewLive Copperway demo · PII vault
  • Compression & PII vaultNewSovereign path controls in Playground
  • Sovereign Exchange5-year cross-org sovereign plan
  • ComputeBare-metal EU inference & training
  • UPDATED on GitHubAGICY AI desktop beta · source
  • Reserve CapacityPre-construction LOI tiers
  • MarketplaceCompute marketplace
  • AI Video SearchNewFind AI videos · avatars · films · ads
  • Sui Agent RailsPrivacy · Walrus · wallet connect
  • GPUs Rent LiveLiveEU partner GPU now · until COD
  • Compute VouchersSovereign compute credits
  • Model LeaderboardFrontier model rankings
  • PricingSubscription tiers
  • Research HubWave-1 publications index
  • Data CentersWorldwide map & tracker
  • Vasilikos Campus42MW sovereign campus briefing
  • Copperway vs gatewaysConcessive pricing & capability evidence
  • OpenRouter alternativesLiteLLM · Portkey · Copperway
  • EU alternative to OpenRouterCLOUD Act / sovereignty buyer guide
  • Copperway vs LLM gatewaysCapability evidence & SRA economics
  • GDPR-compliant AI hostingEU residency & processing path
  • Cerebras vs GroqInference speed & sovereign options
  • CLOUD Act riskUS parented API exposure
  • AI Act Digital OmnibusArticle 50 transparency duties
  • Sovereign cloud truthLabel vs residency — buyer checklist
  • EU AI ActRegulatory mapping & controls
  • AboutProject identity & status
  • GitHubAGICY AI public source · AGiOS-Ai-EU
  • MissionCharter & sovereignty
  • CareersCulture, benefits & hiring ethos
  • Open PositionsEngineering, research & operations roles
  • Ethics & CharterAnti-misconduct & responsible AI
  • Editorial & MethodologySources, claims, corrections
  • Trust CenterSecurity portal · docs · status
  • InvestInstitutional data room & deal flow
  • Living in EUIndividuals & FOs · Class B allocation
  • International investorsPlan B · Greece / Cyprus rails · Class B
  • Equity participation (legacy)CY & GR individual interest · counsel-gated
  • AcademyAI training programs
  • ContactBriefings & inquiries
  • Privacy PolicyGDPR · data processing
  • Terms of ServicePlatform usage terms
  • Cookie PolicyTracking & consent
  • SRA TermsReserve capacity agreement
  • Gateway Pricing DisclaimerCopperway pricing basis
© 2026 AGICY· PROJECT / BRAND OPERATOR · AGICY HOLDINGS LTD — NAME REGISTRATION APPLICATION COMPLETED · AWAITING APPROVAL · CORP DOCS TO FOLLOWDOC: AGICY.AI · REV 2.0 · SOVEREIGN LEDGER
AGICY