This directory contains smaller or efficiency-oriented image classification architectures. Compared with the LargeNetwork directory, these models generally focus more on compactness, mobile deployment, compression, or architectural efficiency.
It also includes a few methodology-oriented folders such as knowledge distillation and deep compression, so this directory is not only “small CNNs” in the narrow sense.
Model families and methods present in the repo:
1_Lenet52_ZFNet3_SqueezeNet4_XNOR-Net5_MobileNet6_ShuffleNet7_KnowledgeDistillation8_DeepCompression9_FractalNet10_MLP-Mixer11_PolyNet12_XceptionNet
Runnable implementations with the usual training files:
1_Lenet52_ZFNet3_SqueezeNet/SqueezeNetNoBypass3_SqueezeNet/SqueezeNetSimpleBypass3_SqueezeNet/SqueezeNetComplexBypass4_XNOR-Net5_MobileNet/MobileNetV15_MobileNet/MobileNetV25_MobileNet/MobileNetV36_ShuffleNet/ShuffleNetV16_ShuffleNet/ShuffleNetV27_KnowledgeDistillation8_DeepCompression12_XceptionNet
Present in the directory but not currently wired up like the runnable folders above:
9_FractalNet10_MLP-Mixer11_PolyNet
Most runnable subprojects follow a repeated scaffold:
config.pydataset.pymodel.pytrain_and_test.pyrun.sh- log directories with histories, TensorBoard outputs, and prediction visualizations
Typical training behavior:
- multi-dataset support
- TensorFlow / Keras training loops
- early stopping
- TensorBoard logging
- reduce-on-plateau learning rate scheduling
- history and prediction plots written to per-run log folders
The common dataset scaffold in this directory supports:
mnistfashion_mnistcifar10cifar100skin_cancercassava_leaf_diseasechest_xraycrop_disease
These are the same broad dataset groups used by much of the larger classification directory, which makes cross-model comparisons easier.
Folder: 1_Lenet5
This is the smallest and most classical CNN in the directory. It is useful as a compact baseline and as the student model in some of the compression-oriented work.
Why it matters:
- historically important early CNN
- simple baseline for small-image tasks
- easy reference point for knowledge distillation and compression experiments
Folder: 2_ZFNet
ZFNet is an AlexNet-era CNN refinement that became known for improved visualization and better understanding of convolutional feature hierarchies.
Why it matters:
- bridges older large-kernel CNNs and later cleaner deep architectures
- useful for comparing early convolution design choices
Folder: 3_SqueezeNet
Implemented variants:
SqueezeNetNoBypassSqueezeNetSimpleBypassSqueezeNetComplexBypass
SqueezeNet focuses on reducing parameter count while maintaining useful performance. The core idea is the Fire module, which uses squeeze and expand stages to limit expensive computation.
Why it matters:
- compact architecture with much smaller parameter footprint
- good case study for model size efficiency
- the bypass variants let you compare connectivity design inside the same family
Folder: 4_XNOR-Net
This is the binary / low-precision efficiency-oriented model in the directory. XNOR-style architectures target much cheaper computation by binarizing activations and/or weights.
Why it matters:
- efficiency-first design
- relevant when deployment cost matters more than maximum accuracy
Folder: 5_MobileNet
Implemented variants:
MobileNetV1MobileNetV2MobileNetV3
These models are designed for efficient classification using depthwise separable convolutions and later architectural refinements.
Why they matter:
- mobile and edge deployment
- excellent compact baselines
- clear evolution from simple efficient CNNs to more optimized modern lightweight networks
Folder: 6_ShuffleNet
Implemented variants:
ShuffleNetV1ShuffleNetV2
ShuffleNet focuses on computational efficiency using grouped operations and channel shuffling.
Why it matters:
- efficiency at low compute budgets
- complements MobileNet by exploring a different lightweight design strategy
Folder: 7_KnowledgeDistillation
This folder is methodological rather than a single architecture. It trains:
- a teacher model
- a smaller student model from scratch
- a student model via distillation
In the current code:
- the teacher is AlexNet-based
- the student is LeNet-based
- distillation uses a
Distillerwrapper with teacher/student losses and temperature scaling
Why it matters:
- shows how to transfer knowledge from a larger model into a smaller one
- useful when you want a compact deployable model with better performance than plain small-model training
Folder: 8_DeepCompression
This folder focuses on model compression ideas rather than only designing a new backbone from scratch.
Why it matters:
- studies reducing model size and cost after training
- fits naturally with the efficiency theme of this directory
Folder: 9_FractalNet
This folder exists, but it is not currently set up in the same runnable pattern as the main implemented subprojects.
Interpret it as:
- present in the repo
- not currently exposed through the same standard training scaffold
Folder: 10_MLP-Mixer
This folder is present, but it does not currently appear to have the same runnable project structure as the other active implementations.
Conceptually, it belongs to the “non-convolutional image classifier” family where token mixing and channel mixing are handled by MLP blocks instead of convolutions.
Folder: 11_PolyNet
This folder is present, but like MLP-Mixer and FractalNet, it is not currently organized as a standard runnable subproject in this tree.
Folder: 12_XceptionNet
This is the most modern fully runnable architecture in the small-network directory. Xception extends depthwise separable convolution ideas into a stronger architecture than earlier lightweight CNNs.
Why it matters:
- efficient convolutional design
- good bridge between compact networks and stronger modern CNN performance
Most runnable small-network trainers follow the same pattern:
- pick a dataset with
--type - pick a GPU with
--gpu - build tf.data datasets
- compile the model with Adam and categorical cross-entropy
- train with early stopping, TensorBoard, and LR reduction
- evaluate on the test set
- save
history.json, prediction grids, and aggregate plots
Representative command:
python train_and_test.py --gpu 0 --type cifar10Examples:
python train_and_test.py --gpu 0 --type mnist
python train_and_test.py --gpu 0 --type fashion_mnist
python train_and_test.py --gpu 0 --type chest_xray
python train_and_test.py --gpu 0 --type crop_disease7_KnowledgeDistillation differs from the others because it trains multiple models in one workflow:
- teacher
- distilled student
- student trained from scratch
This makes it one of the most practically useful folders if your goal is model compression rather than just architecture study.
8_DeepCompression is also more method-oriented than architecture-oriented. It belongs in this directory because its target is efficient classification under tighter resource constraints.
The small-network folders are not perfectly uniform, but the general pattern is:
- image classification with resized inputs
- moderate batch sizes
- Adam optimization
- categorical cross-entropy for supervised classification
Some folders are closer to older-style CNN experiments, while others are clearly motivated by deployment efficiency.
- This directory mixes compact architectures with compression/training methods.
- A few folders are present as placeholders or partial experiments rather than fully runnable implementations.
- Many subfolders contain committed logs, so the repo stores both source code and experiment outputs together.
- The repeated trainer structure makes it easy to compare different small models on the same datasets.
If you want to understand the progression of small and efficient classifiers in this repo:
1_Lenet52_ZFNet3_SqueezeNet4_XNOR-Net5_MobileNet6_ShuffleNet7_KnowledgeDistillation8_DeepCompression12_XceptionNet
Treat 9_FractalNet, 10_MLP-Mixer, and 11_PolyNet as partial/present folders unless you plan to flesh them out further.