Training a coding model to paint watercolours with TRL and OpenEnv

Imported from official source

Research

On 23 August, Surya Narreddi posted a beautiful video of watercolours painted by a language model. The model writes JavaScript through p5.brush, a library that "adds natural drawing tools to p5.js". The video went viral fast, over 1.5M views at the time of writing.

The video came with a blog post explaining the training behind an earlier and narrower stage of the project, close-up flowers rather than the full compositions in the video, sadly without open artifacts yet. His site says a full technical report is coming, so ensure you follow him. The original idea is his, coming from the art and design side, where his skills are way beyond mine. My attempt is on the engineering side, reproducing the recipe in the open with every piece published.

Note: for the context behind the project, told by Surya himself, watch this video of his thesis.

In this article I try and reproduce his idea with TRL and OpenEnv. The reference pool dataset, the RL environment, the training scripts and the trained models, all open.

The whole pipeline runs on Hugging Face, end to end:

  • the RL environment and the scorer model as Spaces
  • the pairwise judge through Inference Providers
  • and every artifact on the Hub, gathered in one collection
  • Once the two Spaces are up, the recipe is one command. Duplicate the environment and the scorer model, set two environment variables for the reward mix, and launch:

    hf jobs uv run train/watercolour_grpo.py --flavor h200 --timeout 48h --secrets HF_TOKEN -- \
      --env-url https://<you>-watercolour-env.hf.space \
      --model Qwen/Qwen3.5-35B-A3B --lora --all-linear --bf16 --gradient-checkpointing \
      --subject 'a peach hibiscus' --references 4 \
      --top-p 0.95 --top-k 20 \
      --lr 5e-5 --lr-scheduler constant_with_warmup --warmup-steps 5 \
      --scale-rewards none \
      --steps 110 --n-episodes 240 --num-generations 8 \
      --per-device-batch-size 1 --gradient-accumulation-steps 8 \
      --max-completion-length 8192 \
      --run-tag my-run --out <you>/watercolour-grpo --push-to-hub
    

    This is an extract. The publication continues at the source.

    Read the original at the source: https://huggingface.co/blog/train-to-paint-with-code

    Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Hugging Face — imported from official source
    Official source
    https://huggingface.co/blog/feed.xml RSS
    Imported
    September 15, 2026 19:08
    Versions
    1 recorded
    Identity
    https://huggingface.co/blog/train-to-paint-with-code

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.