Reference architecture and planning economics

A local runtime you can reason about—and cost.

A transparent starting point for running strong open-weight language models inside an organization's own trust boundary—with architecture, capacity and cost assumptions exposed.

Sovereign boundaryAdaptive routingPermission-aware contextHuman authority

01 / Reference architecture

Local by default. Open at the model boundary.

The durable system is not the GPU server alone. It is the governed path around models, data and human authority.

01

Access & governance

Identity, purpose, policy, quota and approval enter before model execution.

02

PULSAR control plane

Gateway, router, model registry, evaluation gates and release policy stay separate from models.

03

Permitted data plane

Retrieval, object storage and tools inherit source authority and remain network-bounded.

04

Inference plane

Open-weight models run in isolated pools with quantization, batching and task-aware routing.

05

Evidence & operations

Tracing, quality evaluation, power telemetry and human review shape every release.

02 / Capacity scenarios

Start with a workload envelope—not a shopping list

Memory fit depends on model architecture, precision, context length, concurrency and serving software. Benchmark before purchase.

R01

Research node

Evaluation, retrieval, fine-tuning experiments and controlled team use

Compute
2 × 96 GB server/workstation accelerators
Memory
192 GB accelerator · 512 GB ECC system
Storage
8 TB NVMe · protected model registry
Network
25–100 GbE · isolated management
Facility
1.5–2.0 kW design envelope
Quantized 70B-class models, smaller dense models, embeddings and rerankers
REFERENCE CAPEX$50,000
R03

Shared research cluster

Multiple model services, large-context research and capacity expansion

Compute
8 × 96 GB server accelerators
Memory
768 GB accelerator · 2 TB ECC system
Storage
30 TB NVMe · durable shared object storage
Network
Dual 100/200 GbE · separate storage fabric
Facility
7–9 kW design envelope
Larger MoE experiments, multiple resident models and bounded multi-team service
REFERENCE CAPEX$220,000

Reference prices include accelerator hardware, host platform, ECC memory, NVMe storage, network/security components and an integration contingency. They are planning estimates—not vendor quotes.

03 / Türkiye landed cost

Make the assumptions visible

Calculation date

Planning calculator

Edit every rate before a procurement decision.

Base system$110,000
CIF$113,300
Customs duty$0
Import VAT$22,660
Net landed cost, VAT excluded$115,000TRY 5,425,527
Gross cash outlay, VAT included$137,660TRY 6,494,596
CUSTOMS

0% is a planning default—not a tariff ruling

Confirm GTIP classification, origin, trade regime, Incoterm and any additional duty with a licensed customs broker before ordering.

VAT

Cash outlay and economic cost are different views

Import VAT is shown in gross cash outlay. Whether it is recoverable depends on the taxpayer and use; obtain tax advice.

REFRESH

Update the calculation date at least monthly

Refresh accelerator prices, supplier quote, TCMB selling rate, freight, tariff and VAT—and always recalculate on the customs declaration date.

04 / Sustainable operation

Useful intelligence per unit of energy

Sustainability begins with routing and measurement—not a larger server.

Right-size models

Prefer the smallest path that meets the evaluated quality threshold.

Measure real work

Track energy, utilization, queue time and accepted output by workload.

Share capacity safely

Use batching, quotas and isolated pools before adding hardware.

Sources & method

Method: reference CapEx × freight/insurance = CIF; customs duty is applied to CIF; import VAT is applied to CIF plus duty; the clearance allowance is added separately. This simplified public model excludes financing, KKDF, local reseller margin, building works, electricity and recoverability timing.

The thinking behind the system

Why local operation matters

Open-weight models make it possible to inspect the model artifact, choose the serving stack, control the network path and keep sensitive context inside an accountable environment. Local operation is not automatically secure or economical, however. The surrounding architecture determines whether permissions, evidence, updates and human authority survive real use.

Project PULSAR treats hardware as one replaceable layer inside a larger operating model. Identity, policy, routing, permitted context, evaluation, observability and release governance must remain stable while accelerators and models evolve.

An open research baseline

The reference systems on this page are not product bundles or procurement recommendations. They are public research envelopes for comparing:

  • model memory and concurrency needs;
  • latency, throughput and quality under representative workloads;
  • quantization and routing trade-offs;
  • security and operational controls;
  • energy, cooling and utilization;
  • capital cost and Türkiye import cash flow.

The starting assumption uses 96 GB professional server accelerators because this creates a useful, modular memory unit for open-weight inference. Equivalent accelerators and future architectures should be compared against the same workload and governance tests.

How to read the cost model

Reference CapEx is expressed in USD and includes the accelerator set, host platform, ECC memory, local NVMe, network/security components and an integration contingency. It excludes building works, financing, electricity, local reseller margin and contractual support unless stated.

For Türkiye, the calculator separates net landed cost excluding import VAT from gross cash outlay including import VAT. A 0% customs-duty input is only a planning default. Classification, origin and the current import regime must be confirmed before every order.

Prices, tax rates and exchange rates expire. Record the calculation date, refresh the assumptions at least monthly and recalculate using the official exchange rate and tariff position on the customs declaration date.

The decision rule

Do not buy the largest system that fits the budget. Select a small set of representative workloads, establish quality and latency thresholds, measure the smallest viable configuration and preserve an expansion path. Capacity earns its place through accepted output—not theoretical model size.