Frequently Asked Questions
Q: What about context size? What happens after a 30-minute conversation?
With local AI, context size (how much of the chat history the model remembers) eats up RAM fast. On a single, memory-constrained machine, a 30-minute conversation usually means the model runs out of context space and starts “forgetting” the beginning of the chat, or simply crashes.
Because RAMDeck pools the memory of multiple machines, your available context limit scales up with your hardware. If you plug an old 16GB laptop into your cluster, that extra pooled RAM can be dynamically dedicated to expanding your context window. A long conversation that would have choked a single machine can now fit comfortably across the cluster, keeping the model sharp and fully aware of the entire session history.
Q: Won’t using my network as RAM be incredibly slow? A gigabit connection (128MB/s) is slower than a cheap SSD swap file!
Your math on gigabit is right, but it’s not really doing what you’re picturing. Where you’re totally right is actual OS-level swap, where a model doesn’t fit and pages get shuffled to disk constantly — that genuinely is bandwidth-bound and it thrashes hard.
But RAMDeck isn’t swapping. It doesn’t mail model weights back and forth every time someone needs to read something. It hands out different “chapters” (layers) of the model to different machines once, at the start. Each machine keeps its chapters in its own RAM for the whole session.
When generating a response, the machines just pass a short summary note (the activation state) down the relay line to each other. What crosses the network is just that short note, not the weights themselves. Real measurements put it around 650KB per token. Even at a decent generation speed (10 tokens/second), that’s only ~6.5MB/sec — nowhere close to a gigabit network’s ceiling. The real bottleneck is the latency of handing that note to the next machine (which is why wired ethernet beats Wi-Fi for RAMDeck), not the bandwidth.
Q: Since distributing chapters at load time takes 50-60 seconds over the network, wouldn’t it be faster to just keep duplicate copies of the model saved on every device’s local SSD?
You are absolutely correct that loading from local SSDs on every machine would eliminate that initial network transfer time. However, it comes with a massive tradeoff: storage bloat and sync burden.
If we kept a local copy everywhere, every device on your network would need enough free disk space to hold the entire model, not just its assigned piece. If you want to run a 30GB model across a 4-node cluster, you’d be burning 120GB of total storage across your fleet. Furthermore, every time you want to try a new model or update an existing one, you’d have to manually copy it to every single machine again.
By using a single-source-of-truth architecture, the primary machine holds the model on its disk, and RAMDeck pushes the required pieces to the cluster’s RAM only when the model starts up. Yes, you wait ~60 seconds for the network transfer when booting a massive model, but you save massive amounts of local storage across your devices and never have to deal with out-of-sync nodes. And remember, that transfer only happens once at startup—during actual generation, the weights stay completely locked in RAM.
Q: Does any of my data leave my network?
No. RAMDeck runs in offline mode by default — inference calls are restricted to loopback and private-network (RFC1918) addresses unless you explicitly allowlist an external host. Nothing reaches the internet unless you deliberately configure it to.
Q: Is RAMDeck open source?
It’s source-available, not OSI-approved open source, and we say that plainly rather than let people assume otherwise. The public repo runs Apache 2.0 with a Commons Clause addendum — you can read, run, modify, and audit the code freely for personal or internal use, but reselling it as a competing hosted product requires coming to us for a commercial license.
Q: What devices can actually join my cluster?
Windows, macOS, Linux, and Android. Node contribution works across CPU, NVIDIA CUDA, and Apple Silicon Metal backends, and RAMDeck’s cross-platform RPC layer enforces strict build parity so a Windows CUDA machine, a Mac running Metal, and a CPU-only laptop can all work together in the same cluster without mismatch errors.
Q: How much RAM do I actually need to run a given model?
It depends on the model tier. Your catalog spans Starter (Llama 3 8B needs 6GB pooled RAM, 1 node), Pro (Qwen3.6 27B needs 20GB pooled RAM, 2 nodes), up to Max (a 235B-parameter model needs 128GB pooled RAM across 3+ nodes). The dashboard checks this live against your actual connected devices and tells you exactly which models you currently qualify for.
Q: Is my phone safe to use as a node — won’t it kill the battery or overheat?
The Android app only contributes while the phone is charging, connected to Wi-Fi, above a 20% battery floor, and it automatically pauses before thermal throttling or before hitting Android 15’s 6-hour daily foreground-service limit. It’s built to protect your phone first and treat it as a supplemental, not mission-critical, node.
Q: Are the performance numbers real, or marketing figures?
Real, and we ship the benchmark tool specifically so you don’t have to take our word for it. Every run gets logged with a manifest and per-prompt results, and we’ve publicly flagged our own anomalous readings — one early run showed a suspiciously high 1,065 tok/s that the tool itself caught as a cache-hit artifact, not genuine throughput, and we called it out rather than quietly using it in marketing.
Q: Can I plug RAMDeck into tools I already use, like VS Code or my own scripts?
Yes — RAMDeck exposes a standard OpenAI-compatible API (chat completions, streaming, tool/function calling), so anything that already supports a custom base URL — the OpenAI Python SDK, Continue.dev, Cline, or your own code — can point at your cluster with just a URL and API key swap.
Take Back Control of Your AI
Our pre-launch campaign is now live. Join the waitlist on our Indiegogo page to get priority access when we launch.
View Indiegogo Campaign