Agents, coordinators, project structure, VPS access, Forgejo repo, service health checks, Headscale management, contributing guide. |
||
|---|---|---|
| .github | ||
| docs/android-app | ||
| AGENTS.md | ||
| capture_service.py | ||
| config.yaml | ||
| deskcam_context.py | ||
| deskcapture.service | ||
| launcher.py | ||
| LICENSE | ||
| README.md | ||
| reset_db.py | ||
| SPEC.md | ||
| web_app.py | ||
DeskCam — Live Paper/Whiteboard Capture System
DeskCam turns an Android phone (running any MJPEG-capable camera app) into a continuous desk camera. Every few seconds it grabs a frame, runs OCR, and stores the result in a searchable SQLite database. Wren (the AI assistant) reads the latest capture as conversational context without any manual intervention.
Project Status
⚠️ Active development — the Android side (IP Webcam streaming server app) does not yet have a confirmed working open-source replacement. The server-side pipeline (ingestion, OCR, storage, web UI) is fully built and running.
Current architecture: Android phone with "IP Webcam" (Google Play) → MJPEG stream → VPS (capture_service.py) → OCR → SQLite FTS5 → Web UI (port 9001) + Context API.
In progress: Building a custom open-source Android streaming APK from scratch.
Quick Start
Prerequisites
- VPS with Docker/Linux (tested on Hostinger + Debian)
- Python 3.10+, Tesseract OCR, OpenCV
- Android phone with camera (any MJPEG-capable app)
Installation
# Clone the repo
git clone https://forgejo-kkhv.srv1628019.hstgr.cloud/deskcam/deskcam.git
cd deskcam
# Install system dependencies
sudo apt install tesseract tesseract-ocr-eng python3-pip
pip install opencv-python pillow pytesseract peewee flask flask-basicauth pyyaml watchdog
# Configure
cp config.yaml.example config.yaml
# Edit config.yaml — set camera URL, auth, frame interval
# Run
python3 launcher.py
Configuration
Edit config.yaml:
camera:
url: "http://PHONE_IP:PORT/video" # IP Webcam MJPEG endpoint
stream_user: "" # auth user (if set)
stream_pass: "" # auth pass (if set)
frame_interval: 5 # seconds between captures
enabled: true
storage:
data_root: "/data/deskcapture"
thumb_size: [400, 300]
ocr:
lang: "eng"
deskew: true
web:
host: "0.0.0.0"
port: 9001
auth_user: "admin"
auth_pass: "" # set via DESKCAPTURE_AUTH env var
Architecture
┌──────────────────────────────────────────────────────────────┐
│ Android Phone (IP Webcam app — Google Play) │
│ ┌─────────────────────────────────────────────────────────┐│
│ │ MJPEG HTTP stream on port 8080 ││
│ │ http://PHONE_IP:8080/video ││
│ └─────────────────────────────────────────────────────────┘│
└─────────────────────────┬────────────────────────────────────┘
│ HTTP/MJPEG
▼
┌──────────────────────────────────────────────────────────────┐
│ VPS: capture_service.py (background daemon) │
│ - opencv-python reads MJPEG stream frame-by-frame │
│ - Samples a frame every N seconds (configurable) │
│ - Saves original to originals/YYYY-MM-DD/{uuid}.jpg │
│ - Writes to context/latest.json (so Wren can read it) │
└─────────────────────────┬────────────────────────────────────┘
│ frame
▼
┌──────────────────────────────────────────────────────────────┐
│ Frame Processor (same process, same loop) │
│ - Pillow preprocess: deskew, contrast normalization │
│ - Tesseract OCR: extract plain text │
│ - Content type detection: text | diagram | table | mixed │
│ - Summary generation (truncated first line of OCR) │
└─────────────────────────┬────────────────────────────────────┘
│ structured data
▼
┌──────────────────────────────────────────────────────────────┐
│ SQLite + FTS5 (/data/deskcapture/captures.db) │
│ - captures table: uuid, ts, text, text_clean, content_type │
│ - FTS5 virtual table on text_clean for full-text search │
│ - context/latest.json written after each capture │
└─────────────────────────┬────────────────────────────────────┘
│
┌───────────┴────────────┐
▼ ▼
┌────────────────┐ ┌────────────────────────┐
│ Context API │ │ Web UI (Flask) │
│ (JSON file) │ │ port 9001 │
│ │ │ browse/search/live │
└────────────────┘ └────────────────────────┘
Key Files
| File | Purpose |
|---|---|
capture_service.py |
Main daemon — reads MJPEG, processes frames, writes to DB |
web_app.py |
Flask web UI + search API (port 9001) |
deskcam_context.py |
Wren's hook — reads context/latest.json |
launcher.py |
Starts capture_service and web_app as background processes |
config.yaml |
All configuration (camera URL, intervals, auth) |
SPEC.md |
Full technical specification |
deskcapture.service |
systemd unit file for auto-start |
API Endpoints
Context (for AI assistant)
GET /api/context/latest — Returns most recent capture as JSON:
{
"uuid": "abc123",
"ts": "2026-05-24T14:32:01",
"text": "extracted text here",
"content_type": "text",
"summary": "Math notes from CS lecture — derivatives and integrals",
"image_url": "/api/capture/abc123/image"
}
GET /data/deskcapture/context/latest.json — Same data as a file (no HTTP needed)
Web UI
| Route | Description |
|---|---|
/ |
Browse/search captures with thumbnails |
/capture/{uuid} |
Detail view: image + extracted text + metadata |
/live |
Live MJPEG proxy (shows camera feed) |
/api/search?q= |
Full-text search, returns JSON |
Database Schema
CREATE TABLE captures (
id INTEGER PRIMARY KEY AUTOINCREMENT,
ts TIMESTAMP DEFAULT (datetime('now')),
uuid TEXT UNIQUE NOT NULL,
orig_path TEXT NOT NULL, -- path to original image
text TEXT NOT NULL, -- extracted text (raw OCR output)
text_clean TEXT NOT NULL, -- cleaned/normalized text for search
content_type TEXT NOT NULL, -- 'text' | 'diagram' | 'table' | 'mixed' | 'other'
summary TEXT, -- brief human summary
metadata TEXT -- JSON: dims, confidence, etc.
);
CREATE VIRTUAL TABLE captures_fts USING fts5(
text_clean,
content='captures',
content_rowid='id'
);
File storage:
- Originals:
/data/deskcapture/originals/{date}/{uuid}.jpg - Thumbnails:
/data/deskcapture/thumbs/{uuid}.jpg
Systemd Service
sudo cp deskcapture.service /etc/systemd/system/
sudo systemctl enable deskcapture
sudo systemctl start deskcapture
sudo journalctl -u deskcapture -f # watch logs
Contributing / Building the Android App
See docs/android-app/README.md for the in-progress Android streaming APK project.
License
MIT