---
title: Silhouette-based parametric 3D modeling with agents
description: A silhouette-based workflow for agents using the latest AI models makes this possible in practice.
lang: en
articleId: silhouette-modeling
translationKey: silhouette-modeling
version: 1
permalink: /en/articles/silhouette-modeling/
revisionSummary: Initial version.
pubDate: 2026-10-04
tags: [3d-modeling, segmentation, agents]
draft: false
illustrated: true
showDescription: false
contents: true
preserveText: true
---

The latest AI models are very good at 3D modeling. They can model an object from images. With the correct workflow, they can create a parametric model and get the measurements right with only one real-world measurement. A silhouette-based workflow for agents using the latest AI models makes this possible in practice.

In my previous article, I used a structure-from-motion-inspired workflow for agents to model a sewing machine motor from 20 images [@alatalo2026motor]. The method worked great! There is an interactive demo here: <https://janne-alatalo.github.io/agentic-structure-from-motion> The agents detected features in the images and matched them across the images. These features were used to iteratively improve a parametric 3D model and estimate the camera poses and lens projection parameters.

Although the results were amazing, the workflow had room for improvement. It is not very efficient and still needs quite a lot of images, even though it is a clear improvement over standard structure-from-motion workflows.

While modeling that sewing machine motor, I briefly tested the possibility of using silhouette-based optimization for camera pose estimation and object modeling, with promising results. I had the feeling that the silhouette-based method would be more efficient and would work with fewer images.

In the silhouette-based method, the 3D model and camera poses are estimated using silhouettes segmented from the images. A silhouette is the 2D outline of a 3D object projected onto a 2D plane. Figure 1 illustrates this idea. The different 2D planes represent views from different camera angles.

<figure class="article-figure" id="figure-1">
  <a class="figure-image " href="/images/articles/silhouette-modeling/silhouette-geometry-full.webp" aria-label="Open Figure 1 at full resolution">
    <img src="/images/articles/silhouette-modeling/silhouette-geometry-1200.webp" srcset="/images/articles/silhouette-modeling/silhouette-geometry-640.webp 640w, /images/articles/silhouette-modeling/silhouette-geometry-1200.webp 1200w, /images/articles/silhouette-modeling/silhouette-geometry-1920.webp 1920w, /images/articles/silhouette-modeling/silhouette-geometry-full.webp 2800w" sizes="(min-width: 1008px) 960px, (min-width: 900px) calc(100vw - 48px), (min-width: 778px) 730px, calc(100vw - 48px)" width="2800" height="2066" alt="A coffee-table model projected onto three image planes, with a camera and silhouette for each view." loading="lazy" decoding="async" />
  </a>
  <figcaption>Figure 1. Illustration of a 3D object&#x27;s silhouette from three camera angles..</figcaption>
</figure>

The idea of using silhouettes for modeling is not new. Silhouettes are already used to estimate camera poses [@wu2020silhouette] and optimize 3D meshes [@worchel2022mesh]. However, implementing this as an agentic workflow for the new generation of AI models is something that I haven't seen anybody else do yet.

Modern AI agents open up the interesting possibility of improving silhouette-based 3D modeling. They can create an adjustable parametric model of the imaged object and use optimization algorithms to find the best parameter values for the adjustable features. This simplifies the optimization problem compared to mesh-based modeling. There are far fewer parameters for the optimization algorithm to optimize. However, for this to work, the AI model needs to be smart enough to create an initial 3D model that is close enough to the pictured object. Luckily, the most recent AI models have very good world knowledge and spatial understanding and can get the first model close enough to kickstart the optimization.

## Silhouette extraction with image models

When I was writing the previous article, I tested Meta's SAM 3 model for segmenting object silhouettes from the images [@carion2025sam3]. Unfortunately, that model is really dumb. For example, it did not understand what a pulley is. Prompting the model with "large disk" worked for segmenting the pulley from the sewing machine motor, but it was an unfortunate workaround. Getting the model to segment more specific details was impossible.

I still got the workflow working well enough to test the method, and it did work. I was pretty sure that if the segmentation worked better, the workflow would work very well. Unfortunately, with the SAM 3 model, the segmentation performance was too poor for this workflow.

The idea of segmentation-based 3D modeling stayed in the back of my mind. I saw the release blog post for the new Qwen-Image 2.1 model and was very impressed by its image editing performance [@qwen2026image21]. That release blog post also brought back memories of seeing a cool demo by Pedicini [@pedicini2025editing] that illustrates the capability of image models to edit images while keeping other parts of the image exactly the same.

I realized that this feature could be used for segmentation. The image model is prompted to change only part of the image. For example: "Highlight the coffee table in this image in magenta while coloring the clutter in cyan." After modification, a diff is taken between the original and edited images. The result can be used as a segmentation map. Figure 2 illustrates the process using OpenAI's ChatGPT Images 2.5 model [@openai2026images25].

<figure class="article-figure" id="figure-2">
  <a class="figure-image workflow-image" href="/images/articles/silhouette-modeling/segmentation-workflow-full.webp" aria-label="Open Figure 2 at full resolution">
    <img class="workflow-overview" src="/images/articles/silhouette-modeling/segmentation-workflow-1200.webp" srcset="/images/articles/silhouette-modeling/segmentation-workflow-640.webp 640w, /images/articles/silhouette-modeling/segmentation-workflow-1200.webp 1200w, /images/articles/silhouette-modeling/segmentation-workflow-1920.webp 1920w, /images/articles/silhouette-modeling/segmentation-workflow-full.webp 3564w" sizes="(min-width: 1008px) 960px, (min-width: 900px) calc(100vw - 48px), (min-width: 778px) 730px, calc(100vw - 48px)" width="3564" height="1144" alt="Three stages: the original coffee-table photograph, a magenta table and cyan clutter edit, and the resulting segmentation map." loading="lazy" decoding="async" />
    <span class="workflow-panels" aria-hidden="true">
      <span class="workflow-panel panel-1"><img src="/images/articles/silhouette-modeling/segmentation-workflow-full.webp" width="3564" height="1144" alt="Original photograph of a coffee table with a tablecloth and knitting." loading="lazy" decoding="async" /></span>
      <span class="workflow-panel panel-2"><img src="/images/articles/silhouette-modeling/segmentation-workflow-full.webp" width="3564" height="1144" alt="Image-model edit: the table is magenta; the cloth and knitting are cyan." loading="lazy" decoding="async" /></span>
      <span class="workflow-panel panel-3"><img src="/images/articles/silhouette-modeling/segmentation-workflow-full.webp" width="3564" height="1144" alt="Segmentation map: visible table pixels are magenta, occluding objects are cyan, and the background is black." loading="lazy" decoding="async" /></span>
    </span>
  </a>
  <figcaption>Figure 2. Using OpenAI&#x27;s ChatGPT Images 2.5 model to segment a coffee table.</figcaption>
</figure>

Figure 3 shows the segmentation map outlines overlaid on the original image. The result is a very accurate segmentation of the table.

<figure class="article-figure" id="figure-3">
  <a class="figure-image " href="/images/articles/silhouette-modeling/segmentation-outlines-full.webp" aria-label="Open Figure 3 at full resolution">
    <img src="/images/articles/silhouette-modeling/segmentation-outlines-1200.webp" srcset="/images/articles/silhouette-modeling/segmentation-outlines-640.webp 640w, /images/articles/silhouette-modeling/segmentation-outlines-1200.webp 1200w, /images/articles/silhouette-modeling/segmentation-outlines-1920.webp 1920w, /images/articles/silhouette-modeling/segmentation-outlines-full.webp 2640w" sizes="(min-width: 1008px) 960px, (min-width: 900px) calc(100vw - 48px), (min-width: 778px) 730px, calc(100vw - 48px)" width="2640" height="1980" alt="The original coffee-table photograph with a magenta table outline and cyan outlines around the tablecloth and knitting." loading="lazy" decoding="async" />
  </a>
  <figcaption>Figure 3. Original image with segmentation map outlines.</figcaption>
</figure>

These modern image models are very accurate, and they are also very smart. Way smarter than the SAM 3 model. They can be used to segment very specific parts of the image. This is just what my proposed workflow needs. In the example image, the model has identified that the knitting is not part of the table and segmented it as an occluding object. The same goes for the tablecloth.

## The agentic workflow for silhouette-based modeling

I tested OpenAI's GPT-6 Astra and GPT-6.1 Sol for this workflow, and both work. GPT-6.1 Luna is too dumb to understand the current instructions. Clearer or more detailed instructions might make it work, but I have not tested that yet. For now, Astra works perfectly, and my $100/month plan gives me enough usage. Therefore, I do not need to try to optimize the workflow for a smaller model.

The workflow for the coffee table would go like this: The agent sees that the images show a coffee table. The agent creates an adjustable 3D model that resembles the table in the images. In this example, the agent ended up with a model with the parameters listed in Table 1 (the first model was slightly simpler, but the agent improved it during optimization).

<p class="table-caption" id="table-1-caption">Table 1. Optimizable parameters in the final model created by the agent.</p>
<div class="table-scroll parameter-table-scroll" tabindex="0" role="region" aria-label="Table 1, scroll horizontally if needed">
<table class="parameter-table" aria-labelledby="table-1-caption">
  <thead><tr><th scope="col">Table part</th><th scope="col">Parameters</th><th scope="col">Count</th></tr></thead>
  <tbody>
    <tr><td>Overall/tabletop</td><td>Table height, top thickness, underside inset, corner radius</td><td>4</td></tr>
    <tr><td>Legs and layout</td><td>Leg inset, upper width, foot width, taper-start height, foot offset, leg depth adjustment, depth-direction inset adjustment</td><td>7</td></tr>
    <tr><td>Shelf</td><td>Height and thickness</td><td>2</td></tr>
    <tr><td>Apron rails</td><td>Height and thickness</td><td>2</td></tr>
    <tr><td>Total count</td><td></td><td>15</td></tr>
  </tbody>
</table>
</div>

The agent has enough world knowledge to know that, for example, all four table legs are the same, so the model shares parameters across all the legs. In addition to the model parameters, each image has parameters for azimuth, elevation, roll, distance, horizontal framing offset, target height, and focal length. These define the camera pose and projection. Altogether, the scene had 36 parameters: 15 for the geometry and 7 for each of the three images.

The tabletop dimensions were given in the prompt as the only real-world measurements. The prompt for this modeling task was:

> _Model the coffee table in the coffee-table directory. Use the cad skill. The tabletop is 65 cm both ways exact. Aim for 98% IoU for all images. Do not model the crap on the table._

The agent segments the silhouette of the table from each image using the image model and the method described above.

The optimization task is then to maximize the intersection over union (IoU) between the extracted silhouettes and the corresponding silhouettes of the projected 3D model. This requires iteratively adjusting the model parameters, camera poses, and lens projection parameters. The agent is free to choose the optimization method that best suits the setup.

Silhouette-based optimization is not a convex optimization problem. There are local maxima that are not global maxima. This complicates the workflow, but the agent handles these problems perfectly. The agent changes the optimization strategy, fine-tunes the initial conditions, and updates the model on the fly depending on the situation while working towards perfect alignment between the model and the images.

IoU gives a good goal for the agentic loop. You can prompt the model to work until it reaches a specific IoU. This gives the agent a clear target and pushes it to keep working until the goal is reached.

Below is an interactive widget that shows the final result. You can rotate the model to see it from different angles. When you click one of the images, the virtual camera aligns with the photo. Adjust the image opacity or click to hide the image to see the difference between the model and the real image.

<aside class="demo-placeholder">

demo here...

</aside>

The workflow is described as a skill for agents here: <https://github.com/janne-alatalo/image-to-cad-skill>. The skill includes a workbench for inspecting the results. It has been tested with OpenAI models, but could potentially work with any model and agent harness that has access to an image model capable of handling the segmentation workflow. Codex automatically provides agents with tools for prompting OpenAI's image models, so no additional setup is needed. At the time of writing this article, Anthropic doesn't have its own image model. With Anthropic models, some additional setup is needed so that the agent can prompt an external image model.

## Furniture rearrangement app as a use case for this workflow

My partner suggested rearranging the furniture in our apartment. I'm probably the one who will end up dragging them around until she finds an arrangement she likes. I started thinking about how to save myself some effort. With "work smarter, not harder" in mind, I decided to make an app for her to try different arrangements virtually before I have to move any actual furniture.

The workflow works very well for furniture modeling because furniture is usually easy to model with a few parameters (straight edges, many parts share the same parameters, etc.). I've been using the workflow to model all of our furniture. Giving just one real-world measurement for each piece of furniture is enough to get the scale right across the models.

In addition to that, I have the floor plan of my apartment as an image. Modeling the apartment from the floor plan is an easy task for a model like GPT-6 Astra.

Once the 3D models were done, it was easy to vibe code an app where she can plan the new arrangement.

Currently, the app just enables her to drag our furniture around, but an interesting future improvement would be to add a feature for importing new furniture from images. The user could upload a few images of a new piece of furniture, and the backend would use Codex through the Codex App Server [@openaiAppServer] to run an agent with access to the skill. I was thinking that this could be useful when shopping for furniture. Take a few photos in the store and see what the furniture would look like at home before buying it. Modeling a simple piece of furniture takes 10-20 minutes, but that is not too long to hang around in a furniture store.

Overall, this was a successful project. It is cool to find new use cases for these new AI models. These are things that would have been science fiction just a few years ago.
