Roadmap

What’s next for VTM Spark, in the order it ships.

  1. VTM-1.5.1

    Base model

    What VTM Spark runs today.

  2. Next

    VTM MoE 100M

    100M parameters, 30M active

    Trained on significantly more data. It arrives as an experimental build you can try, then becomes the default.

    View more details
    What MoE means
    Mixture of experts. The model is split into many small specialists, and each frame only runs the few it needs. It holds 100M parameters but runs about 30M at a time, so it stays quick on your card.
    A new way to fine-tune
    It is the first model trained with an experimental fine-tuning method, on significantly more data than VTM-1.5.1.
    Just your character
    It learns to draw only your character and leave the background out. No work goes into scenery nobody sees, so more goes into you.
  3. Then

    VTM Studio Rigging

    Platform, with two new models

    A platform for making and preparing your character. It rigs your character automatically, then lets you fix what it got wrong.

    View more details
    Character editing model
    An image generation model for editing your character, run locally on your own graphics card.
    Rigging model
    Built like VTM-1.5, but specialized in poses. It generates the separate images of your character that the VTM models are trained to animate.
    Fix before you go live
    Check the rig and correct what it got wrong. The cleaner the rig, the better the output.
  4. Later

    VTM MoE 200M

    200M parameters

    A larger model trained for better speed and a significant step up in quality, with speculative decoding to make frames faster.

    View more details
    What speculative decoding means
    A fast draft step makes each frame first, extremely quickly. The MoE then takes that draft and improves its quality. The aim is more frames, sooner, at the MoE's quality.
    Better quality
    Twice the parameters of the 100M model, trained to give cleaner, more detailed frames.
    Targeted generation (exploring)
    Teaching the model to focus on the parts of the frame that matter most, so they come out sharper.

Follow along, try the experimental builds and talk to the team on Discord.