Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Put your idle devices to work and become your own AI cloud
Demonstrating an asynchronous LLM queuing system that uses idle devices for batch inference, showcasing batch model evaluation, hardware utilization, and automated model deployment.
I am the founder of Kalavai, a tool that turns any device into a scalable platform for GenAI. It helps developers aggregate compute from any source (cloud, on prem, laptops) in a unified layer, and manages one-click distributed deployment of AI models.
What I’d love to present is our new LLM queuing system, which is an asynchronous batch processing queue that helps developers optimise their workloads at scale, much faster and cost effective than real time inference. Behind the scenes, the queue handles batch requests, whilst workers (any computer) pick inference jobs and run them locally. Workers do just-in-time model deployment, which means there is no model idle time (costly!). And because inference jobs are batched, we can optimise model throughput by configuring batch size.
During the presentation, I plan to show how easy it is to run a batch evaluation of multiple models on a test dataset on my computer. This will demonstrate: 1) the power of personal computing devices 2) how queuing maximises hardware utilisation 3) how easy it is to auto deploy models
Kalavai self-hosts diverse AI models on distributed devices via Docker.
LLM Queue optimizes enterprise AI inference via asynchronous batch processing.
Compose Email
Loading recent emails...