oyd@lab: ~

oyd@lab:~$ ./boot --on-your-device

[ ok ] gpu ....... 4× RTX 3090 · 3× DGX Spark

[ ok ] weights ... on local disk

[ ok ] network ... no cloud required

[ ok ] privacy ... data never leaves your device

On Your Device Super Intelligence

Private AI on your hardware

We choose the hardware, deploy and tune language models that run on your own infrastructure. Your data stays in-house, and there is no per-request bill.

oyd@lab:~$

oyd@lab:~cat services.md

What we do

Order the whole path or just one part of it.

01 · When you plan to buy hardware

Hardware selection and benchmarks

We compare GPU servers, NVIDIA DGX Spark and Apple Silicon machines on your tasks and models. We measure speed, context length and quality, so you buy what you actually need.

  • Speed and quality benchmarks
  • Checks of used hardware
  • A recommendation backed by numbers
02 · When the hardware is already there

Deployment and configuration

We prepare a server or a Mac, install a model server with an OpenAI-compatible API, and set up access, automatic restart and monitoring.

  • exllamav3, vLLM, llama.cpp, MLX
  • OpenAI-compatible API
  • Access keys and monitoring
03 · When the system is too slow or too weak

Model selection and optimization

We pick an open model for your task and get more out of the hardware: quantization, speculative decoding, long context, image understanding.

  • Qwen, DeepSeek, GLM and other open models
  • EXL3 and GGUF quantization
  • Speculative decoding
04 · When the model has to work with your data

Wiring it into work

We connect the local model to what your team already uses: search across company documents, a chat or coding assistant, integrations over an API.

  • Document search (RAG)
  • Chat and coding assistants
  • API integrations

oyd@lab:~oyd bench --list

Benchmarks from our lab

Numbers from our own systems. We run the same kind of measurements with your tasks before proposing a setup.

stdout

> 0 results

> First benchmarks are on the way

> We will publish only what we measure ourselves on our own hardware: the model, the settings and the result. For now, the hardware is ready.

oyd@lab:~lsdev --lab

qtydevicememstatus
4×NVIDIA RTX 309024 GBin use
3×NVIDIA DGX Spark128 GBin use
1×Mac Studio M5 Ultra—arriving
1×NVIDIA RTX 2080 Ti11 GBin use
1×NVIDIA RTX 306012 GBin use
2×AMD RX 4808 GBin use
2×Mac (Apple Silicon)—in use
./all-benchmarks

oyd@lab:~cat process.txt

How we work

  1. [1/4] Tasks and data

    We work out which jobs the AI should do, with what data and under which privacy requirements.

  2. [2/4] Measurements

    We test several models and hardware options on your tasks: speed, quality and context length. The numbers decide.

  3. [3/4] Deployment

    We set up the hardware, models, API and access inside your network. The system starts on its own after a reboot.

  4. [4/4] Upkeep

    We monitor the system and try new models as they come out. We upgrade only when measurements show a gain.

oyd@lab:~man local-ai

When local AI makes sense

+

Data stays in-house

Contracts, customer data or source code stay in your network. No need to negotiate what may be sent to an outside provider.

+

No per-request bill

Costs depend on the hardware and electricity, not on how much your team uses the system.

+

You control the model

The model will not change or disappear without your decision. Swap it when a better one comes out.

!

Not always the best choice

The largest cloud models are still stronger at some tasks. If local does not pay off in your case, we will say so after the measurements.

oyd@lab:~oyd --faq

Frequently asked questions

> What hardware do I need?

It depends on the task and the model size. Smaller models run on a single good GPU or an Apple Silicon machine; larger ones need more memory or several devices. We find out with measurements before you buy.

> Will a local model be as good as ChatGPT?

For some tasks yes, for others no. That is why we start with tests on your real tasks and show you the results before you decide.

> Does the data really stay inside?

The model runs on your hardware, in your network. If any part, such as updates, needs the internet, we document it, and that connection can be switched off.

> Can you work with hardware we already have?

Yes. We measure what it can do and tell you which models and tasks are realistic on it.

oyd@lab:~mail info@webedge.dev

Want to find out whether local AI fits you?

Tell us what work you want to run and what hardware you have or plan to buy. We will suggest what to measure first.

./send-an-email

Vilnius, Lithuania · working across the EU · LT / EN / RU