Level 1
Three fully detected cases with distinct layouts and usable DA3 depth.
FoR-T2I-v1-000565
Create an eye-level scene. A helicopter is present, facing image-right. A vendor is present. The vendor sits on the forward side of helicopter, using helicopter's own orientation. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | vendor vs helicopter: anchor_front | SAM fallback: True




| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| helicopter | owl | 0.7521 | [-1.68682861328125, 629.8218994140625, 1586.420166015625, 1568.8282470703125] | 22.7454 |
| vendor | sam | 0.9220 | [1279.0, 1075.0, 1560.0, 1559.0] | 16.5972 |
FoR-T2I-v1-000410
Create an eye-level scene. A tortoise is present, facing image-right. A horses is present. The horses sits on the left-hand side of tortoise, using tortoise's own orientation. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | horses vs tortoise: anchor_left | SAM fallback: False



| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| tortoise | owl | 0.8528 | [1013.19140625, 943.295654296875, 1737.60498046875, 1280.768310546875] | 10.9717 |
| horses | owl | 0.0671 | [312.1605529785156, 628.8869018554688, 981.58544921875, 1284.481689453125] | 11.1558 |
FoR-T2I-v1-000769
Create an eye-level scene. A small vehicle is present, facing away from the viewer. A cat is present. The cat is behind small vehicle, opposite small vehicle's own facing direction. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | cat vs small vehicle: anchor_back | SAM fallback: False



| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| small vehicle | owl | 0.0529 | [816.9669189453125, 862.3418579101562, 1422.5306396484375, 1349.3448486328125] | 5.5336 |
| cat | owl | 0.1438 | [525.7352294921875, 1103.9869384765625, 840.5181884765625, 1441.0521240234375] | 4.0966 |
Level 2
Three fully detected cases with distinct layouts and usable DA3 depth.
FoR-T2I-v1-000056
Create an eye-level scene. A hedgehog is present, facing away from the viewer. A cow is present. The cow sits on the forward side of hedgehog, using hedgehog's own orientation. A cat is present. A cat is placed in the foreground, closer to the viewer, using the image frame. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | cow vs hedgehog: anchor_front; cat vs hedgehog: ['foreground'] | SAM fallback: True




| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| hedgehog | owl | 0.4979 | [823.9176025390625, 975.5914306640625, 1246.2779541015625, 1508.6605224609375] | 4.6933 |
| cow | sam | 0.9220 | [1135.0, 683.0, 1723.0, 1383.0] | 5.7161 |
| cat | owl | 0.2221 | [263.4246826171875, 714.6163330078125, 822.12451171875, 1665.2584228515625] | 4.1245 |
FoR-T2I-v1-000365
Create an eye-level scene. A bear is present, facing image-left. A cell phone is present. The cell phone sits on the left-hand side of bear, using bear's own orientation. A horse is present. A horse is placed in the background, farther from the viewer, using the image frame. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | cell phone vs bear: anchor_left; horse vs bear: ['background'] | SAM fallback: False



| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| bear | owl | 0.0279 | [164.39114379882812, 912.5760498046875, 1129.541748046875, 1564.2742919921875] | 8.0865 |
| cell phone | owl | 0.0090 | [1106.3416748046875, 950.990966796875, 1442.9049072265625, 1571.0439453125] | 7.8712 |
| horse | owl | 0.2512 | [1072.0059814453125, 691.7738647460938, 1883.6348876953125, 1506.9091796875] | 9.4796 |
FoR-T2I-v1-000794
Create an eye-level scene. A skunk is present, facing image-right. A frog is present. The frog is behind skunk, opposite skunk's own facing direction. A chicken is present. A chicken is placed near the top of the image, using the viewer's image frame. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | frog vs skunk: anchor_back; chicken vs skunk: ['image_top'] | SAM fallback: False



| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| skunk | owl | 0.3523 | [683.3858642578125, 959.3890991210938, 1635.4644775390625, 1347.9493408203125] | 2.9979 |
| frog | owl | 0.2905 | [251.5096435546875, 1150.075927734375, 550.08984375, 1305.40087890625] | 3.0600 |
| chicken | owl | 0.4668 | [886.5270385742188, 309.0177307128906, 1126.6820068359375, 686.923828125] | 4.5537 |
Level 3
Three fully detected cases with distinct layouts and usable DA3 depth.
FoR-T2I-v1-000453
Create an eye-level scene. A shrimp is present, facing image-left. A sea turtle is present. The sea turtle faces the opposite direction from the shrimp. The sea turtle sits on the rear side of shrimp, using shrimp's own orientation. A man is present. The man sits on the forward side of sea turtle, using sea turtle's own orientation. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | man vs sea turtle: anchor_front; sea turtle vs shrimp: anchor_back | SAM fallback: True




| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| shrimp | sam | 0.9179 | [271.0, 974.0, 672.0, 1179.0] | 4.4141 |
| sea turtle | owl | 0.6458 | [745.946533203125, 888.3275146484375, 1371.104736328125, 1219.2076416015625] | 4.8006 |
| man | sam | 0.9649 | [1480.0, 860.0, 1741.0, 1184.0] | 4.6151 |
FoR-T2I-v1-001057
Create an eye-level scene. A brown bear is present, facing the viewer. A ostrich is present. The ostrich faces away from the brown bear. From brown bear's perspective, ostrich is to its left. A dog is present. The dog sits on the forward side of ostrich, using ostrich's own orientation. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
eye_level | dog vs ostrich: anchor_front; ostrich vs brown bear: anchor_left | SAM fallback: False



| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| brown bear | owl | 0.3264 | [38.205291748046875, 737.659423828125, 805.048583984375, 1695.937744140625] | 16.3127 |
| ostrich | owl | 0.7247 | [778.80810546875, 393.85504150390625, 1458.33251953125, 1632.637939453125] | 19.3719 |
| dog | owl | 0.1889 | [1496.650634765625, 993.7987060546875, 1923.205810546875, 1683.7762451171875] | 16.3263 |
FoR-T2I-v1-001161
Create a top-down map. A golf cart is present, facing image-left. A suv is present. The suv faces the golf cart. The suv sits on the rear side of golf cart, using golf cart's own orientation. A cow is present. The cow is on suv's own right side, from suv's perspective. All objects are shown at a similar visual size, with small clear gaps between them so none of the objects touch.
top_down | cow vs suv: anchor_right; suv vs golf cart: anchor_back | SAM fallback: False



| Object | Source | Score | BBox [x1,y1,x2,y2] | Median depth |
|---|---|---|---|---|
| golf cart | owl | 0.3544 | [284.359375, 877.294189453125, 872.232421875, 1267.49072265625] | 9.8729 |
| suv | owl | 0.0846 | [979.8800048828125, 930.0122680664062, 1737.2081298828125, 1272.38330078125] | 9.6024 |
| cow | owl | 0.3007 | [1215.5987548828125, 691.347900390625, 1606.2537841796875, 904.9639892578125] | 10.8094 |