Model Orchestration in C++

Introduction

Welcome back to the third lesson of "Building and Applying Your Neural Network Library"! You've made tremendous progress in this course. In our first lesson, we successfully modularized our core neural network components — dense layers and activation functions. Then, in our previous lesson, we organized our training components by creating dedicated modules for loss functions and optimizers. We now have a well-structured foundation with a clean separation of concerns.

However, as you may have noticed from our previous training examples, we're still writing quite a bit of boilerplate code for each training session. We manually create layers, set up optimizers, write training loops, handle forward and backward passes, and coordinate all these components ourselves. While this gives us complete control, it also means we're repeating the same orchestration logic every time we want to train a model.

In this lesson, we're going to orchestrate all these components into a unified, high-level interface. We'll build a powerful Model abstract class that acts as the conductor of our neural network orchestra, coordinating layers, optimizers, and loss functions through clean, intuitive methods like compile(), fit(), and predict(), as well as a SequentialModel concrete subclass that can be seen as a better and improved replacement for the manual training approach we developed previously, providing a much more elegant and maintainable API for building and training neural networks. Let's get started!

The Need for Orchestration

Think of a symphony orchestra — while each musician is skilled at playing their individual instrument, the magic happens when a conductor coordinates all these talents toward a unified performance. Similarly, we've built excellent individual components (layers, optimizers, losses), but we need a conductor to orchestrate them into a seamless training experience.

Currently, our training process requires us to manually coordinate several moving parts:

  • Instantiate layers and build our network architecture.
  • Create an optimizer with specific parameters.
  • Define our loss function.
  • Implement the training loop with forward passes, loss calculations, backward passes, and weight updates.

This manual orchestration is error-prone and repetitive — exactly the kind of work that should be automated. What we need is a model class that serves as this conductor, providing a high-level API that handles the complexities of training coordination while still giving us the flexibility to customize our network architecture, choose different optimizers and loss functions, and control training parameters.

// Current goal: elegant and extensible training API
SequentialModel model;
model.add(std::make_unique<DenseLayer>(2, 4, "relu", 0.1));
model.add(std::make_unique<DenseLayer>(4, 1, "sigmoid", 0.1));
model.compile("sgd", 0.5, "mse");
model.fit(X_train, y_train, 1000);

Understanding Pure Virtual Functions

Let's start by designing our orchestrator, which is a base Model class that defines the common interface and shared functionality for all types of neural networks. We'll use a pure virtual function approach, which allows us to define the contract that all model types must follow while leaving room for specific implementations.

A pure virtual function in C++ is a special type of virtual function that cannot be implemented in the base class but must be implemented by derived classes. Think of it like an architectural blueprint for houses — you can't live in the blueprint itself, but it defines the essential structure that all houses built from it must have (foundation, walls, roof, etc.). In C++, we declare pure virtual functions using the syntax virtual functionName() = 0;. When we mark methods like _forward() and _backward() as pure virtual, we're saying, "Every model type must have these methods, but each one will implement them differently."

This approach is perfect for our neural network library because different model architectures (sequential, convolutional, recurrent) all need the same core functionality — they all need to perform forward passes, backward passes, and training — but each implements these operations differently. By using pure virtual functions, we enforce a contract that guarantees every model type will have the methods our training system expects, while allowing each model to implement them in the way that makes sense for its architecture. This prevents bugs (you can't forget to implement a required method) and ensures our fit() method will work with any model type we create in the future.

Designing the Base Model Class

Now that we understand the role of pure virtual functions, let's implement our Model orchestrator. We'll start with the header file:

include/neuralnets/models/model.hpp

#pragma once
#include <vector>
#include <memory>
#include <string>
#include <stdexcept>
#include <Eigen/Dense>
#include "neuralnets/layers/dense.hpp"
#include "neuralnets/optimizers/sgd.hpp"
#include "neuralnets/losses/mse.hpp"

namespace neuralnets {

class Model {
protected:
    std::vector<std::unique_ptr<DenseLayer>> layers;  // Layer stack - populated by subclasses
    std::unique_ptr<SGD> optimizer;                   // Training optimizer
    std::function<double(const Eigen::MatrixXd&, const Eigen::MatrixXd&)> loss_fn;
    std::function<Eigen::MatrixXd(const Eigen::MatrixXd&, const Eigen::MatrixXd&)> loss_fn_derivative;
    bool is_compiled;                                 // Compilation status flag

public:
    Model();
    virtual ~Model() = default;
    
    void compile(const std::string& optimizer_name = "sgd", 
                double learning_rate = 0.01, 
                const std::string& loss_name = "mse");
    
    virtual Eigen::MatrixXd _forward(const Eigen::MatrixXd& inputs) = 0;
    virtual void _backward(const Eigen::MatrixXd& d_loss_wrt_prediction) = 0;
    
    Eigen::MatrixXd predict(const Eigen::MatrixXd& X);
    void fit(const Eigen::MatrixXd& X_train, const Eigen::MatrixXd& y_train, 
             int epochs, int batch_size = -1, bool verbose = true);
};

} // namespace neuralnets

Now let's implement the base functionality in the source file:

src/models/model.cpp

#include "neuralnets/models/model.hpp"
#include <iostream>
#include <random>
#include <algorithm>

namespace neuralnets {

Model::Model() : is_compiled(false) {}

void Model::compile(const std::string& optimizer_name, double learning_rate, const std::string& loss_name) {
    // Set up the optimizer - easily extensible for new optimizers
    if (optimizer_name == "sgd") {
        optimizer = std::make_unique<SGD>(learning_rate);
    } else {
        throw std::invalid_argument("Unsupported optimizer: " + optimizer_name);
    }

    // Set up the loss function - easily extensible for new loss functions
    if (loss_name == "mse") {
        loss_fn = mse_loss;
        loss_fn_derivative = mse_loss_derivative;
    } else {
        throw std::invalid_argument("Unsupported loss function: " + loss_name);
    }
    
    is_compiled = true;
    std::cout << "Model compiled with optimizer: " << optimizer_name 
              << " (lr: " << learning_rate << "), loss: " << loss_name << std::endl;
}

} // namespace neuralnets

This foundation establishes all the essential components our model will need to coordinate:

  • Layer stack management: The layers vector holds our network architecture using smart pointers for automatic memory management.
  • Optimizer coordination: Centralized optimizer configuration and management.
  • Loss function setup: Unified loss function and derivative handling using function objects.
  • Compilation validation: The is_compiled flag prevents training misconfiguration.
  • Future extensibility: The string-based configuration system makes adding new optimizers and loss functions straightforward.

Implementing Training and Prediction Logic

Now let's add the core training and prediction functionality. The pure virtual function pattern we're using ensures scalability — any new model architecture can plug into our training orchestration system by simply implementing the required methods.

src/models/model.cpp (continued)

Eigen::MatrixXd Model::predict(const Eigen::MatrixXd& X) {
    return _forward(X);
}

void Model::fit(const Eigen::MatrixXd& X_train, const Eigen::MatrixXd& y_train, 
                int epochs, int batch_size, bool verbose) {
    if (!is_compiled) {
        throw std::runtime_error("Model must be compiled before training. Call model.compile().");
    }
    
    int num_samples = X_train.rows();
    if (batch_size == -1) {
        batch_size = num_samples;  // Full batch gradient descent
    }

    // Random number generator for shuffling
    std::random_device rd;
    std::mt19937 gen(rd());

    for (int epoch = 0; epoch < epochs; ++epoch) {
        double epoch_loss = 0.0;
        
        // Create indices for shuffling
        std::vector<int> indices(num_samples);
        std::iota(indices.begin(), indices.end(), 0);
        std::shuffle(indices.begin(), indices.end(), gen);
        
        // Process data in batches
        for (int i = 0; i < num_samples; i += batch_size) {
            int actual_batch_size = std::min(batch_size, num_samples - i);
            
            // Create batch matrices
            Eigen::MatrixXd X_batch(actual_batch_size, X_train.cols());
            Eigen::MatrixXd y_batch(actual_batch_size, y_train.cols());
            
            for (int j = 0; j < actual_batch_size; ++j) {
                X_batch.row(j) = X_train.row(indices[i + j]);
                y_batch.row(j) = y_train.row(indices[i + j]);
            }

            // Coordinated forward-backward-update cycle
            Eigen::MatrixXd y_pred = _forward(X_batch);                    // Forward pass
            double loss = loss_fn(y_batch, y_pred);                       // Loss calculation
            epoch_loss += loss * actual_batch_size;                       // Accumulate loss
            
            Eigen::MatrixXd d_loss_wrt_pred = loss_fn_derivative(y_batch, y_pred);  // Loss gradient
            _backward(d_loss_wrt_pred);                                    // Backward pass

            // Update all trainable layers
            for (auto& layer : layers) {
                if (layer->d_weights.size() > 0) {
                    optimizer->update(*layer);
                }
            }
        }
        
        // Progress reporting
        epoch_loss /= num_samples;
        if (verbose && ((epochs < 10) || ((epoch + 1) % (epochs / 10) == 0) || epoch == 0)) {
            std::cout << "Epoch " << std::setw(4) << (epoch + 1) << "/" << epochs 
                      << ", Loss: " << std::fixed << std::setprecision(6) << epoch_loss << std::endl;
        }
    }
    
    if (verbose) {
        std::cout << "Training finished." << std::endl;
    }
}

The fit() method orchestrates the entire training process with several key responsibilities:

  • Validation: Ensures the model is properly compiled before training.
  • Batch processing: Handles flexible batch sizes for memory efficiency and training stability.
  • Data management: Shuffles data each epoch to prevent overfitting to data order using C++ standard library random utilities.
  • Training coordination: Manages the forward-backward-update cycle automatically.
  • Progress tracking: Provides informative training progress feedback.
  • Architecture independence: Works with any model architecture through the pure virtual function pattern, with _forward and _backward methods.

This design ensures our training system can scale to handle different model types, training strategies, and dataset sizes without requiring code changes to the core training logic.

Building the Sequential Model

Now we can create our concrete SequentialModel class that implements the pure virtual functions from our base Model. This class represents the familiar stack-of-layers architecture we've been working with, but now with a much cleaner interface and full integration into our extensible framework.

include/neuralnets/models.hpp (addition)

class SequentialModel : public Model {
public:
    SequentialModel() = default;
    SequentialModel(std::vector<std::unique_ptr<DenseLayer>> layers_list);
    
    void add(std::unique_ptr<DenseLayer> layer);
    
    Eigen::MatrixXd _forward(const Eigen::MatrixXd& inputs) override;
    void _backward(const Eigen::MatrixXd& d_loss_wrt_prediction) override;
};

src/models/model.cpp (addition)

SequentialModel::SequentialModel(std::vector<std::unique_ptr<DenseLayer>> layers_list) {
    for (auto& layer : layers_list) {
        add(std::move(layer));
    }
}

void SequentialModel::add(std::unique_ptr<DenseLayer> layer) {
    layers.push_back(std::move(layer));
}

Eigen::MatrixXd SequentialModel::_forward(const Eigen::MatrixXd& inputs) {
    Eigen::MatrixXd current_input = inputs;
    for (auto& layer : layers) {
        current_input = layer->forward(current_input);
    }
    return current_input;
}

void SequentialModel::_backward(const Eigen::MatrixXd& d_loss_wrt_prediction) {
    Eigen::MatrixXd current_d_loss = d_loss_wrt_prediction;
    for (auto it = layers.rbegin(); it != layers.rend(); ++it) {
        current_d_loss = (*it)->backward(current_d_loss);
    }
}

The SequentialModel demonstrates excellent separation of concerns:

  • Initialization flexibility: Supports both empty initialization and pre-built layer lists using move semantics for efficient memory management.
  • Architecture building: The add() method provides the familiar interface for building networks layer by layer using smart pointers.
  • Sequential processing: The forward pass flows data through layers in order, and the backward pass propagates gradients in reverse using reverse iterators.
  • Clean responsibility: Focuses solely on managing a linear sequence of layers while inheriting all training orchestration from the parent class.

This design makes our library highly extensible — we can easily add other model types, like convolutional networks or attention mechanisms, by implementing the same pure virtual interface, while all the training logic remains reusable.

Using Our Orchestrated Model

Now let's see our orchestration in action! Here's how we can solve the XOR problem using our new high-level API, which is much cleaner and more maintainable than our previous manual approach. Notice how this same interface will work seamlessly when we extend our library with new layers, optimizers, or loss functions.

examples/xor_with_model.cpp

#include <iostream>
#include <iomanip>
#include <Eigen/Dense>
#include "neuralnets/models/model.hpp"
#include "neuralnets/layers/dense.hpp"

int main() {
    using namespace neuralnets;
    
    // XOR dataset
    Eigen::MatrixXd X_train(4, 2);
    X_train << 0, 0,
               0, 1,
               1, 0,
               1, 1;
    
    Eigen::MatrixXd y_train(4, 1);
    y_train << 0,
               1,
               1,
               0;

    // Build the model architecture - easily extensible to new layer types
    SequentialModel model;
    model.add(std::make_unique<DenseLayer>(2, 4, "relu", 0.1));
    model.add(std::make_unique<DenseLayer>(4, 1, "sigmoid", 0.1));

    // Configure for training - ready for new optimizers and loss functions
    model.compile("sgd", 0.5, "mse");

    // Train the model - same interface works for any model architecture
    std::cout << "Training model for XOR problem..." << std::endl;
    model.fit(X_train, y_train, 1000, X_train.rows(), true);

    // Test the trained model
    std::cout << "\n--- After Training ---" << std::endl;
    Eigen::MatrixXd y_pred = model.predict(X_train);
    std::cout << "Input  | True Output | Predicted Output | Rounded Prediction" << std::endl;
    for (int i = 0; i < X_train.rows(); ++i) {
        std::cout << "[" << X_train(i, 0) << " " << X_train(i, 1) << "] |"
                  << std::setw(10) << static_cast<int>(y_train(i, 0)) << " |"
                  << std::setw(15) << std::fixed << std::setprecision(4) << y_pred(i, 0) << " |"
                  << std::setw(18) << static_cast<int>(std::round(y_pred(i, 0))) << std::endl;
    }
    
    return 0;
}

Notice how much cleaner this is compared to our previous training scripts! The beauty of this approach is that it maintains the flexibility to customize every aspect of our network while eliminating the repetitive boilerplate code. The extensible design means we can easily experiment with different architectures, optimizers, or hyperparameters without rewriting the training logic each time.

Compilation and Build Setup

To compile our orchestrated model example, we need to update our CMakeLists.txt to include the new model components:

CMakeLists.txt

cmake_minimum_required(VERSION 3.10)
project(NeuralNets)

set(CMAKE_CXX_STANDARD 17)

find_package(Eigen3 REQUIRED)

# Include directories
include_directories(include)

# Library sources
set(SOURCES
    src/layers/dense.cpp
    src/optimizers/sgd.cpp
    src/losses/mse.cpp
    src/models/model.cpp
)

# Create the library
add_library(neuralnets ${SOURCES})
target_link_libraries(neuralnets Eigen3::Eigen)

# XOR with model example
add_executable(xor_with_model examples/xor_with_model.cpp)
target_link_libraries(xor_with_model neuralnets)

Compile and run:

mkdir build && cd build
cmake ..
make
./xor_with_model

Discussing the Output

When we run our orchestrated model, we can see how it successfully learns the XOR problem while providing clear feedback about the training process. The consistent interface and automated orchestration ensure reliable, reproducible results across different model configurations.

Model compiled with optimizer: sgd (lr: 0.5), loss: mse

Training model for XOR problem...
Epoch    1/1000, Loss: 0.249842
Epoch  100/1000, Loss: 0.219619
Epoch  200/1000, Loss: 0.055226
Epoch  300/1000, Loss: 0.014176
Epoch  400/1000, Loss: 0.006978
Epoch  500/1000, Loss: 0.004460
Epoch  600/1000, Loss: 0.003217
Epoch  700/1000, Loss: 0.002481
Epoch  800/1000, Loss: 0.001689
Epoch  900/1000, Loss: 0.001444
Epoch 1000/1000, Loss: 0.001444
Training finished.

--- After Training ---
Input  | True Output | Predicted Output | Rounded Prediction
[0 0] |         0 |          0.0471 |                  0
[0 1] |         1 |          0.9742 |                  1
[1 0] |         1 |          0.9743 |                  1
[1 1] |         0 |          0.0471 |                  0

The results demonstrate excellent performance — our orchestrated model achieves the same learning quality as our previous manual implementations, but with much cleaner, more maintainable code. The loss decreases smoothly from 0.25 to 0.0014, and the final predictions correctly solve the XOR problem with high confidence. More importantly, the training process is now fully automated, consistently implemented, and ready to scale to more complex problems and architectures.

Conclusion

Outstanding work! We've successfully built the orchestration layer for our neural network library, creating a powerful and elegant high-level API that coordinates all our modular components. Our Model and SequentialModel classes demonstrate how good software architecture can transform complex, error-prone manual processes into clean, automated workflows.

The orchestration pattern we've implemented provides simplified user interfaces, reduced code duplication, improved maintainability, enhanced extensibility for future development, and a solid foundation for scaling to production-level neural network applications.

We're now ready for the final step in our journey: putting our complete neural network library to work on a real-world dataset. In our next lesson, we'll apply everything we've built to a real-world regression problem, demonstrating how our library handles practical machine learning problems with multiple features and realistic data challenges.

Sign up

Join the 1M+ learners on CodeSignal

Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal