Skip to content

Latest commit

 

History

353 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CloudAI - Distributed Task Execution System

Orchestrate Docker-based tasks across worker nodes with real-time monitoring and a web dashboard.

Go Version Docker MongoDB


Overview

CloudAI is a distributed computing platform for executing Docker-based workloads across a cluster of worker nodes. Built with Go for high performance.

Complete Documentation | Getting Started Guide


Key Features

  • Interactive CLI - Manage cluster from command-line
  • Web Dashboard - Real-time React UI for monitoring and management
  • Real-Time Telemetry - WebSocket streaming of cluster metrics
  • Docker Native - Run any containerized workload
  • REST & gRPC APIs - Full programmatic access
  • MongoDB Persistence - Task history and results
  • Worker Registration Handshake - Register workers from CLI, then workers self-report resources
  • Task Scheduling - Risk-aware Task Scheduling (RTS) with Round-Robin fallback
  • Adaptive Optimization - AOD module trains scheduling parameters using historical data
  • Task Queuing - Automatic queuing when resources unavailable
  • Task Cancellation - Graceful and forceful termination
  • Resource Tracking - CPU, Memory, Storage
  • File Storage - Secure file upload/download for task outputs
  • JWT Authentication - User registration and login

Architecture

User Interface (CLI/API)
         ↓
    Master Node ---> MongoDB (Persistence)
    (Go + gRPC)       
         ↓
    ┌────┼────┐
    ↓    ↓    ↓
 Worker Worker Worker (Go + Docker)

Components:

  • Master: Task assignment, worker management, telemetry aggregation, AOD training (gRPC: 50051, HTTP: 8080)
  • Worker: Docker execution, heartbeat monitoring (Port 50052+)
  • Web UI: React-based dashboard for monitoring (Port 3001)
  • Database: MongoDB for persistence

Communication:

  • gRPC for Master ↔ Worker
  • HTTP/WebSocket for monitoring and API (Port 8080)
  • MongoDB for data persistence

Quick Start

Prerequisites

  • Go 1.22+
  • Docker (daemon running)
  • Protocol Buffers compiler (protoc)
  • MongoDB (via Docker Compose)
  • Node.js 18+ (for Web UI)
  • Python 3.8+ (for future agent extensibility)

Installation

# Clone repository
git clone https://github.com/Codesmith28/CloudAI.git
cd CloudAI

# Set up Python virtual environment (for future agent extensibility)
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

# One-time setup (generates proto code, creates symlinks, installs deps)
make setup

# Build master and worker
make all

# Install UI dependencies (optional)
cd ui && npm install && cd ..

Run the System

# Terminal 1: Start MongoDB
cd database && docker-compose up -d

# Terminal 2: Start Master (includes Web UI on port 3001)
./runMaster.sh

# Terminal 3: Start Worker  
./runWorker.sh

⚙️ First time setup? See WEBUI_SETUP.md for:

  • Default login credentials (admin@localhost / ChangeMeAdmin123!)
  • How to customize admin user via .env
  • Port conflict troubleshooting
  • Running full campaign workflows with live monitoring

Your First Task

# In master CLI (use worker ID/address shown by runWorker.sh)
master> register <worker_id> <worker_ip:port>
master> workers                              # List workers
master> task hello-world:latest              # Submit task (scheduler picks worker)
master> monitor task-<id>                    # Watch execution
master> list-tasks                           # View all tasks

See docs/GETTING_STARTED.md for detailed walkthrough


Usage Examples

CLI Commands

# Cluster management
master> status                                    # Cluster overview (live)
master> workers                                   # List all workers
master> register worker-3 192.168.1.102:50052    # Manual registration

# Task operations (scheduler picks worker)
master> task python:3.9 -name my-task -cpu_cores 2.0 -mem 4.0
master> dispatch worker-1 ubuntu:latest           # Direct assignment
master> monitor task-<id>                         # Watch logs
master> cancel task-<id>                          # Cancel task
master> queue                                     # View queued tasks
master> list-tasks running                        # Filter by status

# File management
master> files alice                               # List user's files
master> task-files task-<id> alice                # View task files
master> download task-<id> alice ./output         # Download files

Monitoring

# REST API - Telemetry
curl http://localhost:8080/telemetry | jq
curl http://localhost:8080/workers | jq

# REST API - Tasks
curl http://localhost:8080/api/tasks | jq

# WebSocket (real-time)
wscat -c ws://localhost:8080/ws/telemetry

# Submit Task via REST API
curl -X POST http://localhost:8080/api/tasks \
  -H "Content-Type: application/json" \
  -d '{
    "docker_image": "ubuntu:latest",
    "command": "echo hello",
    "cpu_required": 1.0,
    "memory_required": 512.0,
    "tag": "batch-job",
    "k_value": 3
  }'

Authentication

# Register new user
curl -X POST http://localhost:8080/api/auth/register \
  -H "Content-Type: application/json" \
  -d '{"name":"Alice","email":"alice@example.com","password":"securepassword123"}'

# Login (returns JWT token)
curl -X POST http://localhost:8080/api/auth/login \
  -H "Content-Type: application/json" \
  -d '{"email":"alice@example.com","password":"securepassword123"}'

See docs/DOCUMENTATION.md for complete API reference


Docker Testbench (Heterogeneous Workers)

A full automated testbench is available under testbench/, including a master-driven workflow.

Master-driven test commands

Interactive (inside master>):

test list
test run <smoke|reliability|ui-smoke|evidence|full> [-profile <hetero-small|recovery-lab>] [-out <dir>] [-keep-env] [-ui-smoke] [-scheduler <current|RR|RTS>]
test cleanup

Non-interactive:

./masterNode test list
./masterNode test run <smoke|reliability|ui-smoke|evidence|full> [-profile <hetero-small|recovery-lab>] [-out <dir>] [-keep-env] [-ui-smoke] [-scheduler <current|RR|RTS>]
./masterNode test cleanup

Default artifacts for test run land in results/testbench/<timestamp>-<suite>/.

Host-master topology

Use the host-master testbench topology when the master runs on the host:

  • Compose stack: testbench/docker-compose.host-master.yml
  • Prometheus config: testbench/observability/prometheus/prometheus.host-master.yml
  • Host-routable worker registration: WORKER_SPECS=worker-small=host.docker.internal:55052,worker-medium=host.docker.internal:55053,worker-large=host.docker.internal:55054

Quick host-master run:

make testbench-host-up
./runMaster.sh
make testbench-host-register
make testbench-host-suite

One-command integration + benchmark automation

Run the complete Docker-backed gate and evidence benchmark pipeline:

make testbench-integration

This command runs Go unit-test preflight plus smoke, reliability, ui-smoke, and evidence suites and stores artifacts in results/testbench/<timestamp>-integration/.

CI equivalent: .github/workflows/testbench-integration.yml (manual dispatch + nightly schedule, artifact upload included).

Detailed runbook: docs/TESTBENCH_RUNBOOK.md.

Evidence Benchmark Campaign

Run a multi-scenario evidence benchmark across schedulers and failure modes:

make campaign              # Run smoke benchmark (heterogeneous-smoke workload)
make campaign-full         # Run full campaign (all workloads + all scenarios)

The campaign exercises:

  • Schedulers: RR (Round-Robin), RTS (Risk-aware Task Scheduling), PPO (offline-trained + online learning)
  • Scenarios: baseline, burst, overload
  • Observability: Exports Prometheus metrics, task telemetry, and scheduler diagnostics to results/campaign/

See docs/TESTBENCH_RUNBOOK.md for campaign command-line options.

Recovery Semantics

CloudAI tracks logical tasks separately from physical execution attempts, enabling automatic recovery when workers fail.

  • every worker assignment carries attempt_id and attempt_no
  • if a worker stops heartbeating, the active attempt is marked lost and the logical task is requeued automatically
  • late results from older attempts are recorded for audit but cannot overwrite the current task outcome

Inspection endpoints:

curl http://localhost:8080/api/tasks/<task_id> | jq
curl http://localhost:8080/api/tasks/<task_id>/attempts | jq

Documentation


Troubleshooting

Common Issues

Issue Solution
Worker not connecting Check netstat -tuln | grep 50051, verify firewall
Task fails Run docker pull <image> to test, check logs with monitor
MongoDB error Run docker-compose ps in database/, restart if needed
Authentication error Check JWT_SECRET env, ensure token is valid

Debug Logging:

export LOG_LEVEL=debug
./masterNode

See docs/DOCUMENTATION.md Section 12 for detailed troubleshooting


Contributing

Contributions welcome! Areas of interest:

  • New scheduling algorithms
  • Dashboard/UI implementation
  • Authentication & authorization
  • Performance optimizations
  • Documentation improvements

Process: Fork → Feature Branch → Commit → Push → Pull Request


License

Apache 2.0


Star this repo if you find it useful!

Built with gRPC • MongoDB • Docker

About

Distributed cluster computing platform for executing Docker-based workloads across a cluster of worker nodes with Agentic Scheduling.

Resources

Stars

8 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages