Managed IntelligenceWhat Is a Managed Intelligence Provider?Managed Intelligence ServicesPrivate AI: Your Hardware or AzureOntology & Semantic ModelManaged Agent HarnessAI Cost Calculator
IT ServicesManaged IT ServicesCybersecurityCloud ComputingMicrosoft Copilot
IndustriesFinanceHealthcareLegalEducationManufacturing
AboutOur ApproachCareers
ResourcesLocal vs Cloud AI Cost GuideCustomer ZeroAll Resources
Blog
Contact
Free AI Session

Private AI: Your Hardware or Azure

Private AI on hardware you own or in Microsoft Azure

Models run on hardware you own, typically an NVIDIA DGX Spark, or in your Microsoft Azure subscription on Microsoft Foundry. Your data stays under your control, your costs stay predictable, and Tekscape manages the whole stack end to end.

Private AI means the models that read your data run where you control them: on hardware you own, or in your own Microsoft Azure subscription on Microsoft Foundry. A prompt leaves only when a routing policy you approve allows it. Tekscape, a Managed Intelligence Provider (MIP), sizes, deploys and runs it for you end to end, and sends hard reasoning to frontier cloud models when that is the better deal.

Private AI: models and data inside a secure boundary, with only selected prompts leaving through a gate

Definition

What private AI means

Private AI is AI that runs where you control it. The model sits on a machine in your office or rack, or in a Microsoft Azure subscription that belongs to you. It reads your documents, tickets and records there, and by default nothing is sent to a public AI service.

Local AI is the on-premises version, where the hardware is physically yours. Most often that is an NVIDIA DGX Spark, but we are hardware-agnostic and will run on the machine that fits your workload. Microsoft Azure is the cloud version: the models run in your own tenant on Microsoft Foundry, under your identity, network and access rules. In both cases you decide which prompts, if any, may go to another cloud model.

That is how it differs from a typical AI subscription. With a public API, every prompt and every attached file leaves your network. With private AI, a prompt leaving is the exception, and each one is logged.

Tekscape is a Managed Intelligence Provider (MIP). We do not sell you a box and walk away. We run the models, the agents and the guardrails around them as a managed service, grounded in a semantic model of how your business works.

Deployment

Two ways to run it: your hardware or Microsoft Azure

Both options get the same harness, the same semantic model and the same managed service. You can also mix them: steady, sensitive work on your hardware, and the rest in Azure.

  • On hardware you own

    Typically an NVIDIA DGX Spark in your office or rack. We are hardware-agnostic, so if another workstation or GPU server fits your workload better, we run on that. Best when data must stay in your building or volume is steady and high.

  • In Microsoft Azure

    Your agents run in your own Azure subscription on Microsoft Foundry. There is no hardware to buy. Best when you already run on Microsoft 365 and Azure, or your volume is low or bursty.

Data boundary

What stays inside your boundary, and what leaves

Your data, documents and models sit inside the boundary, on your hardware or in your Azure subscription, and agents query them there.

A routing policy you approve decides which prompts may cross the boundary to a frontier cloud model. The policy blocks the data classes you mark as sensitive from leaving, so the decision is made by rule rather than case by case.

Each call that crosses the boundary is logged: what went out, to which model, and why.

Your office
  • Where it runs

    Your hardware in your office or rack, or your own Azure subscription.

  • Models

    Open-weight models on your hardware, or models you deploy in Microsoft Foundry.

  • Data

    Files, records and the semantic model. Leaves only by a rule you approve.

Policy gate

Checks each prompt against your rules and strips what must stay private.

Frontier cloud model

Used only for the hard cases, when it is the better deal.

Everything inside the dashed line stays on hardware you own or in your own Azure subscription. The only way out is through the policy gate, and only for prompts your rules allow.

Only prompts your routing policy allows cross the boundary. The rest are answered inside it.

Hardware

Hardware tiers: pick the class, not the price tag

We size hardware to your workloads, the number of people using the agents and how much data the agents must read. We scope each deployment against the three classes below.

For scale, one independent arXiv preprint on on-premises LLM cost reports that medium open-weight models run on two 80GB data-center GPUs with under 10% accuracy loss relative to large open-weight models. The same paper notes that commercial APIs still hold a slight edge in peak accuracy. That is one paper's estimate, and your own workloads may behave differently.

Hardware classes for local AI deployment. Prices are not listed because they move; sizing is quoted per engagement.
TierHardware classSuited forModel size
Workstation-classAn NVIDIA DGX Spark, a Mac Studio-class machine or a single high-end GPU workstationA team or department: drafting, summarizing, document Q&A, meeting notes, first-line triageSmall open-weight models; some medium models
Small GPU serverA rack or tower server with one or two data-center GPUsCompany-wide agents used by several people at once; finance, service desk and document workflowsMedium open-weight models
Multi-GPU serverA server with several data-center GPUs, or a small clusterMany agents running at once, larger context windows, heavier reasoning and batch jobs over large archivesLarger open-weight models, or several models side by side

This page lists no prices. Hardware prices move, so we quote sizing per engagement, based on your volume and your data.

Managed end to end

What Tekscape manages

Whether it runs on your hardware or in Azure, we run it. One accountable team covers the full lifecycle.

  • Sizing

    We measure your workloads and volume, then recommend the smallest hardware class that does the job.

  • Install

    We rack, cable, harden and network the hardware on your site, or set up the Azure subscription, network and Microsoft Foundry project in your tenant for a cloud deployment.

  • Model serving

    We deploy and tune the open-weight models and the serving stack your agents call.

  • Monitoring

    24/7 monitoring of uptime, response times, capacity and errors, with alerts to our team.

  • Patching

    We patch the operating system, drivers and serving software on a schedule and test each update before rollout.

  • Model updates

    When a better open-weight model ships, we test it against your tasks before we swap it in.

  • Backups

    We back up configuration, prompts, agent memory and indexes, and we test restores.

  • Security

    Access control, network segmentation, encryption and audit logs around every model and agent.

Models

Open-weight models: capable, and yours to run

Open-weight models publish their weights, so they can run on your own hardware without per-token API charges. Check each model's license terms before use. The stronger models now handle drafting, summarizing, extraction and classification well.

Most business tasks are narrow. You rarely need the largest model to sort an invoice or summarize a ticket. Gartner predicts that by 2027 organizations will use small, task-specific AI models at least three times more than general-purpose LLMs.

We choose models per task, not per brand. Each one is tested on your own examples before it goes live, and tested again when a new version ships.

Hybrid routing

Your deployment by default, frontier models for hard reasoning

Private AI does not mean giving up the strongest cloud models. It means using them on purpose.

Research supports routing. In the RouteLLM benchmark from LMSYS, routing between GPT-4 and a cheaper model cut cost by over 85% on MT Bench while keeping 95% of GPT-4 quality. That benchmark routed between API models. It is not a local-hardware measurement or a customer result, but it supports the design idea of sending each request to the cheapest model that can handle it.

  1. Classify the request

    Before any model sees the request, the harness checks the task type and the data it touches.

  2. Answer inside your boundary when it can

    Routine and sensitive work goes to the open-weight model on your hardware.

  3. Escalate by policy

    Hard reasoning goes to a frontier cloud model, but only when your routing policy allows that data class to leave.

  4. Log and review

    Each escalation is logged. We review the mix and move work to local or to the cloud as costs and models change.

Cloud prices keep falling. a16z estimates that API prices for equivalent capability fall about 10x per year, and Epoch AI measures 9x to 900x per year depending on the benchmark. Both track API prices, not total cost of ownership. We re-check the routing mix as prices move.

Honest fit

When private AI is not the right call

Local AI does not always save money. These are the cases where we will tell you to stay in the cloud.

One independent arXiv preprint puts on-premises LLM break-even at a few months for small models, about 2 years for medium models and about 5 years for large models. It finds on-premises deployment viable mainly at roughly 50M+ tokens per month or under strict data-residency requirements. The paper's cost model counts only GPU cost and electricity and leaves out staffing and maintenance. That gap is where a managed service comes in, and we include staffing and maintenance costs when we run your numbers.

Situations where cloud AI is usually the better deal.
Your situationOur advice
Low volume: a handful of users and occasional promptsStay on a cloud API or an existing subscription. Hardware is unlikely to pay back.
No data constraints: nothing regulated, privileged or confidentialCloud is simpler. Revisit if volume grows or your data changes.
Most tasks need frontier-level reasoningRun cloud-first with routing. Keep local capacity small, or skip it.
A short-term project or a trialProve the use case in the cloud first. Buy hardware once the workload is known.
You want the lowest possible price per token todayCloud prices fall quickly. Local AI buys control and predictability, not an assured discount.

Compliance

Support for HIPAA, legal privilege and financial data obligations

Private AI helps you meet your obligations. It does not make you compliant on its own.

Healthcare: keeping patient data on hardware you control, with access logs, supports your HIPAA safeguards. You still need your own policies, agreements and risk assessments.

Legal: keeping privileged matter files off public AI services helps protect attorney-client privilege and client confidentiality.

Financial data: account, payroll and deal data can stay on-site, with an audit trail of which agent read what, and when.

We document the data boundary, the routing policy and the audit trail so your compliance and legal teams can review them. Your counsel and auditors make the final call.

No tool makes a business compliant by itself. We build controls that support your compliance program.

Next step

Run your own numbers first

Start with the Local vs Cloud AI Cost Calculator. Enter your token volume, API price and hardware estimate to see where break-even lands, with falling cloud prices modeled in.

Then read the honest break-even guide, When does local AI actually save money? It covers the evidence, the counter-arguments and where hybrid routing wins.

If the numbers point to private AI, we size it with you and quote it per engagement.

FAQ

Frequently asked questions

What is private AI?
Private AI is AI that runs on hardware you own or in your own Microsoft Azure subscription, so your data is read there instead of being sent to a public AI service. A prompt leaves only when a routing policy you approve allows it. Tekscape deploys, runs and monitors it for you as a managed service.
Is local AI cheaper than cloud AI?
Sometimes. It depends on your volume, model size and data constraints. One independent arXiv preprint puts break-even at a few months for small models, about 2 years for medium models and about 5 years for large models, and finds on-premises deployment viable mainly at roughly 50M+ tokens per month or under strict data-residency rules. That paper counts GPU cost and electricity only. Use our calculator to test your own numbers.
What hardware do we need for on-premises AI?
We typically start with an NVIDIA DGX Spark, a compact workstation-class machine, but we are hardware-agnostic. We size against three classes: a workstation-class machine, a small GPU server or a multi-GPU server, depending on how many people use the agents, how much data they read and which models the tasks need. If you would rather not own hardware, we run the same agents in Microsoft Azure. Because hardware prices move, we quote sizing per engagement.
Can we run private AI in Microsoft Azure instead of on our own hardware?
Yes. We deploy the same agents in your own Microsoft Azure subscription on Microsoft Foundry, under your identity and access rules, with the same harness and audit logs. It suits businesses already standardized on Microsoft 365 and Azure, or with low or bursty volume. As a Microsoft Cloud Solution Provider, we manage the Azure side too.
Does private AI make us HIPAA compliant?
No tool makes you compliant on its own, but private AI supports your HIPAA safeguards. Keeping patient data on hardware you control, with access logs and a documented data boundary, strengthens your program. Your policies, agreements, risk assessments and counsel still determine compliance.
Can we still use frontier cloud models?
Yes. Private AI routes hard reasoning to frontier cloud models when your policy allows it. Routine and sensitive work stays in your deployment by default. Each escalation is logged, and we adjust the local-versus-cloud mix as prices and models change.
What does Tekscape manage after the install?
We manage model serving, 24/7 monitoring, patching, model updates, backups and security. On your hardware or in Azure, we operate it. We test new open-weight models against your tasks before swapping them in, and we report on usage and routing.

Sources

  1. arXiv 2509.18101: on-premises LLM deployment cost and break-even (preprint)Source for the break-even by model size, the 50M+ tokens per month threshold, and the finding that medium open-weight models run on two 80GB GPUs with under 10% accuracy loss relative to large open-weight models. Preprint, not peer reviewed. Its cost model counts GPU cost and electricity only.
  2. LMSYS: RouteLLMCost cut of over 85% on MT Bench at 95% of GPT-4 quality. The benchmark routed between API models; it is not a local-hardware or customer result.
  3. a16z: LLMflationAPI prices for equivalent capability fall about 10x per year. Covers API prices, not total cost of ownership.
  4. Epoch AI: LLM inference price trends9x to 900x per year depending on the benchmark. Covers API prices, not total cost of ownership.
  5. Gartner prediction on small task-specific models (via CXOToday)By 2027, at least three times more use than general-purpose LLMs.

Find out if private AI fits your business

Bring your volume, your data constraints and your current AI spend. We will give you a straight answer on whether local AI, cloud or a hybrid is the better deal, and size it if it fits.

Talk to Tekscape