ClickClickClick: The Secret Framework Making Computers Run Themselves

B
Bright Coding
Auteur
ClickClickClick: The Secret Framework Making Computers Run Themselves

What if your computer could actually understand what you want—and just do it? Not with brittle scripts that break when a button moves two pixels. Not with RPA tools that cost thousands and require certification courses. I'm talking about genuine autonomous control: you type "draft an email to Rob about lunch," and your machine opens Gmail, composes the message, and saves it as a draft. No hands. No macros. Just pure AI-driven execution.

Sound like science fiction? It's not. It's happening right now with ClickClickClick, and developers are losing their minds over it.

Here's the painful truth we've all accepted: automating computer interactions is broken. Selenium scripts shatter when websites update. PyAutoGUI coordinates fail across screen resolutions. RPA platforms lock you into expensive enterprise contracts. We've been duct-taping solutions together for decades, praying nothing changes.

But what if you could leverage the reasoning power of any LLM—local or remote—to see your screen, understand your goal, and execute the steps? That's exactly what ClickClickClick delivers. Built by the team at InstaVM, this framework is already turning heads with demos that feel like magic: drafting emails, navigating Google Maps, even playing chess on Lichess. And the best part? It's open-source, MIT-licensed, and ready for you to hack on today.

Ready to see how it works? Let's dive in.

What is ClickClickClick?

ClickClickClick is an open-source framework designed to enable autonomous Android and computer control using any Large Language Model—whether you're running Llama 3.2-vision locally through Ollama, hitting OpenAI's GPT-4o API, or leveraging Google's Gemini Flash-Lite.

Created by InstaVM, the project represents a fundamental shift in how we think about UI automation. Instead of hardcoding selectors, coordinates, or DOM paths, ClickClickClick uses a vision-language approach: the LLM looks at your screen, reasons about what it sees, and decides what to click, type, or swipe next.

The framework is built around three core components:

  • Planner: The strategic brain that breaks high-level goals into actionable steps
  • Finder: The visual processor that locates UI elements on screen
  • Executor: The hands that perform clicks, swipes, and keystrokes

This separation of concerns is genius. Each component can use a different model optimized for its specific job. The current codebase is explicitly marked as "highly experimental," which means early adopters have a rare opportunity to shape its evolution. The repository is evolving rapidly with community contributions, and the InstaVM team is actively iterating based on real-world usage.

Why is it trending now? Three forces converged: multimodal LLMs finally became capable enough to parse screenshots reliably, local model performance crossed a usability threshold (qwen3.5:4b and Llama 3.2-vision), and the developer community is starving for alternatives to bloated RPA stacks. ClickClickClick arrived at exactly the right moment.

Key Features That Separate It From the Pack

Let's get technical about what makes this framework special:

🔍 True Vision-Based Interaction Unlike traditional automation that relies on accessibility trees or DOM inspection, ClickClickClick captures screenshots and feeds them to vision-capable LLMs. This means it works on any application—web, native, Android, macOS—without special hooks or SDK integrations. If a human can see it, the framework can interact with it.

🧠 Dual-Model Architecture The Planner/Finder split isn't just architectural elegance—it's a performance optimization. You can pair a fast, cheap model for planning with a powerful vision model for element detection. The README explicitly notes that Gemini 3.1 Flash-Lite currently delivers the best results when used for both roles, but the flexibility to mix and match is where the real power lies.

📱 Cross-Platform by Design Android via ADB. macOS via native automation. The --platform flag switches contexts instantly. One codebase controls your phone and your laptop.

⚡ Four Interface Options Whether you want a quick CLI one-liner, a Gradio web UI for demos, direct Python↗ Bright Coding Blog API integration, or a full REST API for microservices—ClickClickClick meets you where you are. This isn't a toy; it's built for production pipelines.

🖼️ Adaptive Image Quality The --image-quality parameter (1-100) lets you trade resolution for speed. Running local models? Drop to 45% quality for faster inference without sacrificing accuracy. This kind of operational tuning is what separates prototypes from production tools.

🔒 Local-First Friendly Full Ollama support means you can automate sensitive workflows without ever sending screenshots to external APIs. Your data stays on your hardware. For privacy-conscious organizations, this is a massive differentiator.

Real-World Use Cases Where ClickClickClick Dominates

1. Cross-Platform Testing Without Test Scripts

Traditional E2E testing requires maintaining brittle selectors. ClickClickClick enables semantic testing: "Verify that users can complete checkout with PayPal." The LLM figures out the clicks. When the UI changes, it adapts. No more 3 AM pages because a data-testid attribute disappeared.

2. Accessibility Automation for Legacy Systems

That internal tool from 2008 with no API? The vendor portal that only works in IE mode? ClickClickClick doesn't care. If it renders pixels, it can interact. Organizations sitting on decades of technical debt finally have a migration path that doesn't require rewriting everything.

3. Mobile Workflow Orchestration

The Android demos are particularly compelling. "Start a 3+2 game on Lichess" isn't a simple deep link—it's navigating the app, finding the right buttons, selecting time controls. Imagine automating fleet device management, content publishing across dozens of phones, or QA testing on real hardware without Appium's overhead.

4. Personal Productivity Agents

The Gmail draft demo reveals the bigger picture: this isn't just about clicking—it's about task completion. "Find bus stops in Alanson, MI" requires opening Maps, searching, parsing results, and extracting structured information. ClickClickClick bridges the gap between LLM reasoning and real-world action, making it a foundation for genuine personal AI assistants.

5. RPA Replacement for Cost-Conscious Teams

Enterprise RPA licenses can hit six figures annually. ClickClickClick runs on a $20/month API budget—or entirely free on local hardware. For startups and mid-market companies, this cost arbitrage is transformational.

Step-by-Step Installation & Setup Guide

Ready to run your first autonomous task? Here's the complete setup:

Prerequisites

Before starting, ensure you have:

  • Python 3.8+ installed
  • ADB (Android Debug Bridge) for Android automation—this is mandatory even if you only plan to use macOS, as the project structure expects it
  • Git for cloning

Installation

Clone the repository and enter the project directory:

git clone https://github.com/instavm/clickclickclick
cd clickclickclick

Create and activate a virtual environment (strongly recommended):

python3 -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

Install dependencies:

pip install -r requirements.txt

For CLI tool installation (alternative method):

pip install <repo-whl>  # Replace with actual wheel file when available

Initial Configuration

The framework uses a YAML-based model configuration. Create your settings at config/models.yaml:

# Example structure - refer to repository for current format
planner:
  default: gemini
  gemini:
    api_key: ${GEMINI_API_KEY}
    model: gemini-3.1-flash-lite
  
finder:
  default: gemini
  gemini:
    api_key: ${GEMINI_API_KEY}
    model: gemini-3.1-flash-lite
  
  ollama:
    base_url: http://localhost:11434
    model: llama3.2-vision

Export your API keys as environment variables:

export GEMINI_API_KEY="your-key-here"
export OPENAI_API_KEY="your-key-here"  # If using OpenAI models

First-Time Setup

Run the interactive setup to configure your default models:

python main.py setup

This prompts you to select planner and finder models, then validates your API connectivity.

Starting the REST API Server

For API-based usage, launch the FastAPI server:

uvicorn api:app --reload

The server runs on http://localhost:8000 by default.

REAL Code Examples from the Repository

Let's examine actual code patterns from the ClickClickClick repository, with detailed explanations of how each works.

Example 1: Basic CLI Task Execution

The simplest entry point—running a task from command line:

./click3 run "open google.com in browser"

This one-liner demonstrates the framework's core philosophy: natural language as the API. No XPath, no CSS selectors, no coordinate maps. The string "open google.com in browser" goes to the Planner, which decomposes it into steps: (1) locate browser application, (2) activate it, (3) focus address bar, (4) type URL, (5) submit.

The Finder then processes screenshots to locate each element, and the Executor performs the actual interactions. The ./click3 wrapper handles environment setup and model initialization automatically.

Example 2: Platform-Specific Execution with Model Overrides

Here's a more complex invocation showing full parameter control:

python main.py run "Open Google news" --platform=android --planner-model=openai --finder-model=gemini

Breaking this down:

  • --platform=android: Routes execution through ADB to a connected Android device
  • --planner-model=openai: Uses GPT-4o for strategic planning (strong reasoning, higher cost)
  • --finder-model=gemini: Uses Gemini for visual element detection (excellent vision capabilities, cost-effective)

This model mixing is where power users optimize. OpenAI's models excel at complex multi-step reasoning; Gemini's vision models offer superior UI element detection at lower latency. By splitting roles, you get the best of both worlds.

Example 3: Local Model with Performance Tuning

For privacy-sensitive or cost-constrained scenarios:

python main.py run "Open Reddit" --platform=android --planner-model=ollama --finder-model=gemini --image-quality=45

Critical optimizations here:

  • --planner-model=ollama: Runs Llama 3.2-vision or qwen3.5:4b locally—zero API costs, data never leaves your machine
  • --finder-model=gemini: Still uses cloud vision for reliable element detection (local vision models remain weaker at this task per the README's own benchmarks)
  • --image-quality=45: Reduces screenshot size by 55%, dramatically speeding up local model inference

The README explicitly notes that qwen3.5:4b "works as planner (slow, basic navigation) but not reliable as finder." This honest documentation helps you make informed trade-offs. The 45% quality setting is empirically derived—sharp enough for UI element detection, compressed enough for local model throughput.

Example 4: REST API Integration

For production service integration, here's the complete curl pattern:

curl -X POST "http://localhost:8000/execute" \
  -H "Content-Type: application/json" \
  -d '{
    "task_prompt": "Open uber app",
    "platform": "android",
    "planner_model": "gemini",
    "finder_model": "openai"
  }'

Expected response structure:

{
  "result": {
    "status": "success",
    "data": {
      // Execution details: steps taken, screenshots, final state
    }
  }
}

This REST interface transforms ClickClickClick from a local tool into a service-oriented automation backbone. Microservices can queue tasks, webhooks can trigger workflows, and orchestration layers like Airflow or Temporal can manage complex multi-step processes.

The API validates all parameters—invalid platforms or models return 400 Bad Request with descriptive errors. Runtime failures surface as 500 responses with execution traces for debugging.

Example 5: Minimal API Call (Defaults in Action)

curl -X POST "http://127.0.0.1:8000/execute" \
  -H "Content-Type: application/json" \
  -d '{
    "task_prompt": "Open Safari"
  }'

Notice what's absent: no platform specified (defaults to android), no model overrides (planner defaults to openai, finder to gemini), no quality tuning (defaults to 100%). This demonstrates sensible defaults that let you iterate quickly, then optimize once you have working patterns.

Advanced Usage & Best Practices

Optimize Your Model Pairing The README's finding is definitive: Gemini 3.1 Flash-Lite as both planner and finder currently wins. But costs and latency matter. For high-volume operations, consider Ollama planner + Gemini finder. Monitor your specific workloads—task complexity dramatically shifts the optimal pairing.

Master Image Quality Tuning Don't accept defaults blindly. Start at 100% for debugging, then systematically reduce. The README documents 45% as viable for Ollama; experiment with 30-60% for cloud models. Each percentage point directly impacts token costs and response times.

Structure Tasks for LLM Success Vague prompts fail. "Do something with email" confuses the planner. "Open Gmail, click Compose, enter rob@gmail.com in To field, write subject 'Lunch Saturday?', write body paragraph about baby congratulations and Saturday 1PM availability, save as draft" succeeds. Specificity in language translates to precision in execution.

Implement Retry Logic This is "highly experimental" software. Screenshots capture transient states. Network hiccups interrupt API calls. Wrap ClickClickClick calls in exponential backoff, especially for API usage. Log full execution traces for debugging.

Contribute Back The project actively welcomes contributions. Run pre-commit install before developing:

pre-commit install
pre-commit autoupdate
pre-commit run --all-files

Comparison with Alternatives

Feature ClickClickClick Selenium/Playwright PyAutoGUI Enterprise RPA
Setup Complexity Low (pip install) Medium Low High (weeks)
UI Change Resilience High (vision-based) Low (selectors break) None (coordinates) Medium (recordings)
Cross-Platform Android + macOS Web only Desktop only Varies by vendor
Natural Language Input Native None None Limited (NLP addons)
Local Execution Full Ollama support N/A N/A Rare/expensive
Cost Free + API usage Free Free $10K-$100K+/year
Learning Curve Low Medium Low High (certification)
Maturity Experimental Production Stable Enterprise

The Verdict: ClickClickClick occupies a unique position. It's more resilient than coordinate-based tools, more flexible than web-only frameworks, and infinitely more accessible than enterprise RPA. The trade-off is maturity—this is bleeding-edge software for developers comfortable with iteration.

FAQ

Q: Is ClickClickClick production-ready? The README explicitly states "highly experimental." Use it for prototyping, personal automation, and non-critical workflows. Implement robust error handling and never use it for financial transactions without human verification.

Q: Can I run this entirely offline? Yes—with limitations. Ollama supports local Llama 3.2-vision and qwen3.5:4b. However, the README notes local models underperform as finders. For best results, you'll want at least one cloud vision model.

Q: Does it work on Windows or Linux? Currently Android and macOS (osx) are explicitly supported. Windows and Linux support likely requires community contribution—check the repository for latest platform updates.

Q: How much does it cost to run? Free if fully local (hardware costs only). With cloud APIs, a typical task costs $0.01-$0.10 depending on model choice and screenshot complexity. Gemini Flash-Lite is aggressively priced for this use case.

Q: What's the difference between planner and finder? The planner strategizes: "To open Uber, I need to unlock phone, find app icon, tap it, wait for load." The finder executes visually: "Given this screenshot, the Uber icon is at coordinates (342, 891)." Separation allows optimization of each role.

Q: Can it handle multi-step workflows? Absolutely. The demos include drafting emails (multiple fields, composition, saving) and chess game initiation (navigation, time control selection, game start). Complex workflows are its strength.

Q: How do I contribute? Open issues, submit PRs, and follow the pre-commit workflow. The three-component architecture (Planner/Finder/Executor) offers clear extension points for new models, platforms, and execution strategies.

Conclusion

ClickClickClick represents something rare in developer tools: a genuine paradigm shift. We've spent decades teaching machines our languages—XPath, selectors, coordinates—when we could have been teaching them to see and reason like us.

This framework isn't perfect. It's experimental. It requires iteration. But it points toward a future where automation is semantic, not syntactic—where you describe what you want, not how to click it.

For developers building the next generation of AI agents, for teams drowning in legacy system maintenance, for anyone who's ever screamed at a brittle Selenium test at 2 AM—this is your invitation to explore something genuinely new.

Star ClickClickClick on GitHub, clone it, break it, improve it. The repository is waiting. Your computer is ready to start listening.

What will you automate first?


Made with ❤️ by InstaVM | Follow them for updates on this rapidly evolving project.

Commentaires 0

Aucun commentaire pour l'instant. Soyez le premier à réagir !

Laisser un commentaire