Scene Classification with Graph Neural Networks
Classifying six scene categories from Intel Image Classification with a GCN over superpixel graphs, comparing three ways of turning an image into a graph.
- PyTorch Geometric
- Python
- SLIC
- Delaunay
A university group project: classify six scene categories (buildings, forest, glacier, mountain, sea, street) from the Intel Image Classification dataset without a CNN. Every image becomes a graph before it reaches a Graph Convolutional Network. My part was the data pipeline, the models, training and the visualisations.
Why a graph at all#
A CNN assumes an image is an even grid of pixels whose neighbourhoods never change. I wanted to try the opposite: group pixels into homogeneous regions with SLIC, treat each region as a node, and let the relationships between regions become the edges. Framed that way, the real question was never "can a GNN beat a CNN" but how much does the way you build the graph change the answer.
Three graphs from the same image#
I built three variants, one notebook each, over the same 14,034 training and 3,000 test images:
gcn-dt.ipynb: resize to 32×32, one node per pixel, RGB as the feature, edges from a Delaunay triangulation over the pixel coordinates.gcn-combine.ipynb: SLIC cuts the image into roughly 100 superpixels, each node carries mean colour plus centroid (5 dimensions), and Delaunay now runs over the region centroids.gcn_slic.ipynb: SLIC with 400 superpixels on a 100×100 image, edges from region adjacency — two superpixels that touch get an edge — and node features widened to 9 dimensions.
The second variant is the tidiest of the three, because Delaunay hands back an edge list directly without walking the image again:
def create_combined_graph(image, n_segments=100, compactness=10):
image = resize(image, (100, 100), anti_aliasing=True)
segments = slic(image, n_segments=n_segments, compactness=compactness, sigma=1)
regions = regionprops(segments + 1)
# One node per superpixel: its centroid and the mean colour of the region.
centers = np.array([r.centroid for r in regions])
colors = np.array([image[segments == i].mean(axis=0) for i in np.unique(segments)])
# Delaunay over the centroids; each triangle yields three edges, deduped via a set.
edges = set()
for simplex in Delaunay(centers).simplices:
for i in range(3):
edges.add(tuple(sorted((simplex[i], simplex[(i + 1) % 3]))))
x = torch.tensor(np.hstack([colors, centers]), dtype=torch.float)
edge_index = torch.tensor(np.array(list(edges)).T, dtype=torch.long)
return Data(x=x, edge_index=edge_index), segmentsNode features did the heavy lifting#
The first two variants describe a node by colour and position only, and both plateau around 60%. In the last variant I added the shape of the region: eccentricity, aspect ratio, solidity, centroid and perimeter alongside mean colour — nine dimensions in total. That single change moved the score more than anything else, which makes sense: forest, sea and street differ in the shape of their patches, not only in their colour.
The architecture stayed deliberately plain so the comparison meant something: three
GCNConv layers at 128 → 64 → 32 with BatchNorm, a global mean pool, then two Linear
layers with 0.3 dropout.
class GCNModel(torch.nn.Module):
def __init__(self, input_dim, hidden_dims, num_classes, dropout=0.3):
super().__init__()
self.conv1 = GCNConv(input_dim, hidden_dims[0])
self.bn1 = BatchNorm(hidden_dims[0])
self.conv2 = GCNConv(hidden_dims[0], hidden_dims[1])
self.bn2 = BatchNorm(hidden_dims[1])
self.conv3 = GCNConv(hidden_dims[1], hidden_dims[2])
self.bn3 = BatchNorm(hidden_dims[2])
self.fc1 = Linear(hidden_dims[2], hidden_dims[3])
self.fc2 = Linear(hidden_dims[3], num_classes)
self.relu = ReLU()
self.dropout = Dropout(dropout)
def forward(self, data):
x, edge_index, batch = data.x, data.edge_index, data.batch
x = self.relu(self.bn1(self.conv1(x, edge_index)))
x = self.relu(self.bn2(self.conv2(x, edge_index)))
x = self.relu(self.bn3(self.conv3(x, edge_index)))
# Every image has a different superpixel count; mean pool fixes the width.
x = global_mean_pool(x, batch)
return self.fc2(self.dropout(self.relu(self.fc1(x))))Preprocessing split away from training#
Region adjacency has to scan every pixel to find out which superpixels touch. Across
14,034 images that step costs far more than training does, and at first I was paying it
again on every model tweak. So I cut the pipeline in two: convert images to graphs
once, save each graph as a .pt file under graphs/train_graphs/, and skip files
that already exist on a rerun. Training then only loads .pt files into a DataLoader,
so changing the model no longer means rebuilding the dataset.
Outcome#
- The superpixel + region adjacency + 9-feature variant reached 78.17% accuracy on the 3,000 test images, stopping early at epoch 87 with a patience of 10
- The three graph constructions separate clearly: pixel + Delaunay 59.6%, superpixel + Delaunay 62.4%, superpixel + region adjacency 78.2%
buildingsstayed the weakest class throughout: 0.85 precision but only 0.47 recall, mostly leaking intostreet— two scenes with near-identical region texture- All three models ship as a Streamlit app that takes an upload or an image URL and returns the predicted label