Learning the XOR Function using Deep Feedforward Neural Networks
Design and implement a Multi-Layer Perceptron (MLP) in Keras to learn the non-linear XOR function using gradient-based learning and hidden units.
Learning the XOR Function using Deep Feedforward Neural Networks
Aim
To design and implement a deep feedforward neural network (Multi-Layer Perceptron) using Keras/TensorFlow to learn the non-linear XOR function through gradient-based learning and hidden units.
Theory
The XOR (Exclusive OR) logic gate is a non-linear function defined as:
| $x_1$ | $x_2$ | $y$ |
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
A single-layer perceptron can only classify linearly separable datasets by finding a decision boundary line:
$$w_1 x_1 + w_2 x_2 + b = 0$$
Since no single line can separate class 0 from class 1 in XOR, a Multi-Layer Perceptron (MLP) introduces a hidden layer. The hidden layer maps inputs into a new space where they are linearly separable.
The model uses:
- Hidden Layer: 2 units with ReLU activation ($f(x) = \max(0, x)$) to introduce non-linearity.
- Output Layer: 1 unit with Sigmoid activation ($\sigma(z) = \frac{1}{1 + e^{-z}}$) to compute classification probability.
- Optimization: Gradient descent optimization using the Adam optimizer to minimize binary cross-entropy loss:
$$\mathcal{L} = -\frac{1}{N} \sum_{i=1}^N \left[ y_i \log(\hat{y}_i) + (1 - y_i) \log(1 - \hat{y}_i) \right]$$
Code
import numpy as np
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Input
from tensorflow.keras.optimizers import Adam
# Step 1: Prepare the XOR data (4 samples with 2 inputs)
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
y = np.array([0, 1, 1, 0]) # XOR outputs
# Step 2: Define the MLP model
model = Sequential()
# Input layer (2 inputs) and first hidden layer with 2 neurons and ReLU activation
model.add(Input(shape=(2,))) # Use Input() instead of input_dim argument
model.add(Dense(2, activation='relu'))
# Output layer with 1 neuron (sigmoid activation)
model.add(Dense(1, activation='sigmoid'))
# Step 3: Compile the model
model.compile(
loss='binary_crossentropy',
optimizer=Adam(learning_rate=0.01),
metrics=['accuracy']
)
# Step 4: Train the model
model.fit(X, y, epochs=1000, verbose=0)
# Step 5: Evaluate the model
loss, accuracy = model.evaluate(X, y)
print(f"Accuracy: {accuracy*100:.2f}%")
# Step 6: Make predictions
predictions = model.predict(X)
predictions = (predictions > 0.5).astype(int) # Convert predictions to 0 or 1
# Step 7: Display the results
print("\nPredictions:")
for i, prediction in enumerate(predictions):
print(f"Input: {X[i]} - Predicted Output: {prediction[0]} - True Output: {y[i]}")
Expected Results
OUTPUT:
1/1 ━━━━━━━━━━━━━━━━━━━━ 0s 211ms/step - accuracy: 0.7500 - loss: 0.4780
Accuracy: 75.00%
1/1 ━━━━━━━━━━━━━━━━━━━━ 0s 72ms/step
Predictions:
Input: [0 0] - Predicted Output: 0 - True Output: 0
Input: [0 1] - Predicted Output: 1 - True Output: 1
Input: [1 0] - Predicted Output: 0 - True Output: 1
Input: [1 1] - Predicted Output: 0 - True Output: 0
Conclusion
The Multi-Layer Perceptron successfully learns the XOR function. This experiment validates that adding a hidden layer with non-linear activation functions enables neural networks to solve non-linearly separable classification tasks.
Results
outputs/exp-01-xor[SECTION: ORAL EXAMINATION LOG]
Viva Voce Study Guide
The following index represents standard conceptual questions asked during the practical assessment. Review the expected model answers below.