This directory contains larger image classification architectures implemented in TensorFlow. The collection is organized roughly as a historical progression from early deep CNNs to residual, attention, and transformer-based models.
The codebase is best understood as a model zoo of self-contained experiments rather than a single unified package.
Implemented families:
1_AlexNet2_VGGNet3_NetworkInNetwork4_InceptionNet5_ResNet6_HighwayNet7_DenseNet8_ResidualAttentionNet9_SENet10_ResNext11_CapsuleNetwork12_VisionTransformer
Runnable implementations with the usual config.py, dataset.py, train_and_test.py, and often run.sh:
1_AlexNet2_VGGNet/VGG112_VGGNet/VGG11_LRN2_VGGNet/VGG132_VGGNet/VGG16C2_VGGNet/VGG16D2_VGGNet/VGG193_NetworkInNetwork4_InceptionNet/InceptionV14_InceptionNet/InceptionV24_InceptionNet/InceptionV34_InceptionNet/InceptionV45_ResNet/ResNet185_ResNet/ResNet345_ResNet/ResNet505_ResNet/ResNet1015_ResNet/ResNet1526_HighwayNet7_DenseNet/DenseNet1217_DenseNet/DenseNet1697_DenseNet/DenseNet2017_DenseNet/DenseNet2648_ResidualAttentionNet/ResidualAttentionNet568_ResidualAttentionNet/ResidualAttentionNet929_SENet10_ResNext12_VisionTransformer
Present but not structured like the other runnable subprojects:
11_CapsuleNetwork
Most runnable folders follow the same pattern:
config.pydefines input size, batch size, epochs, and learning ratedataset.pybuilds tf.data pipelines for multiple datasetsmodel.pydefines the architecturetrain_and_test.pyhandles training, validation, evaluation, plotting, and loggingrun.shis often included for convenience
Typical outputs:
- TensorBoard logs
history.json- prediction grids such as
predictions.png - aggregate accuracy plots across datasets
- generated architecture diagrams from
plot_model
The common dataset loaders in this directory support:
mnistfashion_mnistcifar10cifar100skin_cancercassava_leaf_diseasechest_xraycrop_disease
In practice, the image-based medical/agriculture datasets are read from local dataset roots, while MNIST/CIFAR are loaded via tf.keras.datasets.
Folder: 1_AlexNet
A classic early deep CNN. Large kernels and stacked convolution blocks make it a useful historical starting point for modern image classification.
Good for:
- understanding the jump from shallow CNNs to large-scale deep vision models
- comparing older convolution design against later residual and efficient models
Folder: 2_VGGNet
Implemented variants:
VGG11VGG11_LRNVGG13VGG16CVGG16DVGG19
VGG-style models rely on repeated small 3x3 convolutions and very uniform stage design. They are heavy in parameter count but conceptually simple.
Why they matter:
- easy to reason about
- strong baseline for “plain deep CNN” design
- useful for comparing depth scaling without residual connections
Folder: 3_NetworkInNetwork
This architecture replaces simple linear convolutional filters with small learned subnetworks, often through 1x1 convolutions.
Why it matters:
- pushes feature abstraction deeper inside each block
- an important step toward later bottleneck and inception-style designs
Folder: 4_InceptionNet
Implemented variants:
InceptionV1InceptionV2InceptionV3InceptionV4
These models use parallel branches within a block to capture multiple receptive-field scales at once.
Why they matter:
- more expressive multi-scale feature extraction
- an important branch in CNN design before residual networks became dominant
Folder: 5_ResNet
Implemented variants:
ResNet18ResNet34ResNet50ResNet101ResNet152
ResNet introduces residual skip connections, allowing much deeper optimization than plain stacked CNNs.
Why they matter:
- foundational modern vision architecture
- strong baseline for deeper supervised classification
- provides the template for many later families
Folder: 6_HighwayNet
Highway networks use learned gates to regulate how much transformed information versus carried-forward information passes through each block.
Why they matter:
- historically important precursor to residual connections
- useful for understanding gated depth before ResNet became the standard
Folder: 7_DenseNet
Implemented variants:
DenseNet121DenseNet169DenseNet201DenseNet264
DenseNet connects each block to all later blocks in the same dense stage, encouraging feature reuse and stronger gradient flow.
Why they matter:
- efficient feature reuse
- improved gradient propagation
- reduced need to relearn similar low-level features repeatedly
Folder: 8_ResidualAttentionNet
Implemented variants:
ResidualAttentionNet56ResidualAttentionNet92
These models combine residual backbones with attention modules that emphasize informative spatial or feature responses.
Why they matter:
- introduces explicit attention into CNN classification
- bridges standard residual models and more structured attention-based networks
Folder: 9_SENet
Squeeze-and-Excitation networks recalibrate channel responses by learning which feature channels should be emphasized or suppressed.
Why they matter:
- lightweight performance-oriented channel attention
- widely reused in later CNN families
Folder: 10_ResNext
ResNext extends the residual idea with grouped transformations, often described through cardinality rather than just depth or width.
Why they matter:
- stronger representational power without only scaling depth
- practical middle ground between plain residual blocks and more elaborate branch-heavy modules
Folder: 11_CapsuleNetwork
This folder exists in the repo, but it does not currently follow the same runnable structure as the other large-network implementations.
Interpret it as:
- present in the repo
- not yet integrated into the same training scaffold as the others
Folder: 12_VisionTransformer
This is the transformer-based image classifier in the directory. Instead of relying purely on convolutions, it treats the image as a sequence of patch embeddings and applies transformer layers for classification.
Why it matters:
- marks the shift from CNN-dominant classification to transformer-based vision models
- useful for comparing convolutional inductive bias against patch-token modeling
Most large-network trainers follow this pattern:
- choose a dataset with
--type - choose a GPU with
--gpu - build tf.data pipelines
- compile the selected model with Adam and categorical cross-entropy
- train with early stopping, TensorBoard, and learning-rate reduction
- evaluate on the test set
- save history and visualization artifacts
Representative command pattern:
python train_and_test.py --gpu 0 --type cifar10Common dataset examples:
python train_and_test.py --gpu 0 --type mnist
python train_and_test.py --gpu 0 --type cifar100
python train_and_test.py --gpu 0 --type chest_xray
python train_and_test.py --gpu 0 --type skin_cancerMost of the classical CNN folders use roughly:
INPUT_SIZE = [224, 224, 3]BATCH_SIZE = 64EPOCHS = 10LEARNING_RATE = 1e-4
Some later folders differ:
SENetandResNextuse64x64input and longer training defaultsVisionTransformeruses224x224but a separate config profile- a few subfolders contain more experimental or inconsistent config values
- The directory mixes polished runnable subprojects with a few partial or experimental folders.
- Logging artifacts are committed in many subdirectories, so these folders contain both code and experiment history.
- Dataset path handling is not perfectly uniform across all subfolders.
- The training scripts are very similar across many models, which makes the directory easy to extend but also means there is some duplication.
If you want to understand the progression of large image classifiers in this repo:
1_AlexNet2_VGGNet3_NetworkInNetwork4_InceptionNet5_ResNet6_HighwayNet7_DenseNet8_ResidualAttentionNet9_SENet10_ResNext12_VisionTransformer
11_CapsuleNetwork should be treated separately because it is not wired into the same common training structure.