Core Features
- Unified Memory Pooling Turn the disparate hardware sitting around your house into a single, cohesive local AI cluster. RAMDeck intelligently shards massive AI models across the combined RAM and VRAM of your existing laptops, desktops, and phones.
- OpenAI-Compatible API Endpoint Zero vendor lock-in and zero friction. RAMDeck exposes a standard OpenAI-compatible `/v1/chat/completions` API endpoint to your local network. Instantly point your existing workflows, scripts, or apps (like LangChain, Continue.dev, or Cline) to your cluster just by changing the URL.
- Cross-Platform RPC Parity Windows CUDA machines, Apple Silicon Macs, plain old CPU-only boxes, and Android devices — RAMDeck bridges all of them into one cluster with strictly enforced build parity. No manual builds, and no driver mismatch drops.
- Live Topology Dashboard Watch your cluster assemble in real-time. The web dashboard provides live visibility into your connected nodes, their memory capacities, active shard assignments, and lets you seamlessly swap the “primary” model host on the fly.
- One-Click Benchmarking Stop guessing if your hardware combination will work. Every RAMDeck cluster ships with its own benchmark button. One click tests your active topology and returns a real, timestamped tokens-per-second (t/s) result.
Community Requested Features
- Enterprise-Grade Knowledge Base (KB) Ingestion You don’t just want to talk to a model; you want to talk to your data. RAMDeck includes a unified upload endpoint that natively supports drag-and-drop ingestion of PDFs, JSONL, and ZIP files. The engine automatically chunks the raw text and coordinates cross-node cache fetching.
- Training Orchestration (Fine-Tuning) RAMDeck isn’t just for running models—it can train them. The hub natively ingests standard Axolotl and LLaMA-Factory YAML configurations. It translates those specs into a cluster job, routing the heavy lifting to your CUDA nodes while bypassing weak CPUs.
The Secret Sauce: Reliability
- Dynamic Context Sizing Context memory scales dynamically with your hardware. Plug in more machines, and that extra pooled RAM unlocks the model’s massive native context window (often 32,000+ tokens) which would normally crash a single machine. If your model’s native context is still too large for your current pool, RAMDeck intelligently scales it down to the maximum safe capacity.
- Persistent Topology State Node configurations and toggles are durably stored. If the central daemon restarts or your power flickers, the hub perfectly recovers your active cluster layout instantly.
- Zombie Node Expiration Advanced liveness checks distinguish between genuine network drops and active multi-gigabyte tensor transfer spikes. If a node is just busy transferring a 9GB layer, RAMDeck won’t accidentally kill it.
- Vulkan Suppression Intelligent service wrapping safely bypasses unsupported integrated graphics hangs on CPU-only nodes, meaning old hardware doesn’t crash the cluster.
Take Back Control of Your AI
Our pre-launch campaign is now live. Join the waitlist on our Indiegogo page to get priority access when we launch.
View Indiegogo Campaign