Computer Vision & Image Processing
A vision-system result becomes comparable when it can be traced from the output, through its geometry or data structure and the operation that produced it, back to the source image and capture condition.
Scene
What enters the system
- capture device
- source scene
- image sequence
- sensor parameters
Representation
How the scene is encoded
- pixels
- regions
- features
- masks
- scene graph
Computation
What changes the representation
- model
- pipeline
- loss
- matching
- decoder
Evidence
What makes behavior observable
- paired outputs
- metrics
- curves
- UI states
input → output
Visual pair
Processing result and the condition under which it was produced.
region / mask / mesh
Spatial object
Coordinate frame, membership rule, and relationship to the source image.
layers + objective
Model record
Data flow through the model and the operation performed at each boundary.
state / timeline
Runtime record
Trigger, transition, intermediate result, and visible system response.
Source
A prompt, reference image, class label, or control signal enters with a stated source and format.
Encode
A conditioning encoder produces a representation; its injection point and interface are identified.
Transform
The representation conditions an iterative denoising process in pixel or latent space.
Decode
A final representation is decoded or rescaled into an output image.
Compare
The output is tied to a test configuration, baseline, metric definition, and paired visual result.
The trace separates the public diffusion substrate from the particular conditioning, sampling, representation, or data-construction mechanism under review.
What does an overlap or detection-loss value actually measure?
IoU and generalized IoU depend on four visible geometric objects: the predicted region, reference region, their intersection or union, and—when used—the smallest enclosing region.
- coordinate convention
- region-construction rule
- formula-to-figure variable mapping
When does a mask, scene graph, feature map, or box become usable evidence?
The object needs defined fields or membership rules, the operation that creates it, and the downstream operation that consumes it. A free-standing visualization leaves the technical relation implicit.
- object schema
- producer
- consumer
- coordinate frame or value range
Where does the technical difference sit: data, training, model, or runtime?
Dataset construction, objective and parameter update belong to training; thresholds, matching, non-maximum suppression, decoding, and other output rules belong to inference. They are recorded on separate paths.
- training-data construction
- objective
- updated parameter
- runtime decision rule
The useful unit of comparison is a transformation at a defined representation boundary. A model name or output image alone does not locate the technical difference.
Intermediate data objects expose what one operation produces and the next consumes.
Training and inference contain different operations and require separate records.
A metric is interpretable only when its geometry, variables, baseline, and test condition are fixed.
- Source and capture conditions are comparable.
- Training records are not substituted for runtime records.
- The same metric definition and baseline are used across the comparison.