The AI Generalist Program : How to find use cases as an AI Generalist?

I feel AI is one of the fastest-progressing fields. Every other year, we see a new model that outperforms the older models. In this lesson, we’ll cover the mathematical concepts behind GANs (Generative Adversarial Networks), Diffusion Models, and VLA (Vision-Language-Action) models in layperson’s terms. Understanding the mathematical foundations helps you see the models differently and find pinpoint use cases.

Generative Adversarial Network: Let’s use a football match as an example. In a football match, every player tries to maximise their score. We see a goal when this converges to a solution. When there is a tie, we have no solution. This concept is from Game Theory. In a Generative Adversarial Network, we have two Neural Networks. One is called the Generator, and the other one is called the Discriminator. Mathematically, both models are differentiable. Why do the neural Networks need to be differentiable? If you remember your Class 12th Mathematics, you will remember finding maxima and minima in Calculus. You are given a function, find its derivative, and then find where it equals 0. Then we take the derivative again and find its direction. If the slope is 0 and the angle is downward, we call it a minimum. Many optimisation algorithms also work similarly. In training a Neural Network, we use a method called Backpropagation. Backpropagation finds the partial derivative at every layer and tries to find the point where the residual error is minimised (residual error is lowest when the model produces outputs close to the training data). For Backpropagation to work, the neural network must be differentiable. Returning to the concept of Generator (G) and Discriminator (D) – they are Neural Networks – the key objective is to train G so D’s mistake-making probability is maximised. G generates an image from noise, and D classifies it as fake or real. Train G in such a way that D classifies G’s generated image (fake) as a real image.

Hands-on Assignment:

Step 1: You can ask Copilot/any other tool you use to generate the script with the following prompt:

Generate a complete Python script to generate an image using a Generative Adversarial Network. Explain the training procedure and comment the script properly. Use the MNIST dataset for this. In addition, visualise how the image is generated from noise.

Step 2: Review the generated script, run it, and observe the outputs.

Diffusion Model: ChatGPT also used this to generate the images. The concept is that, unlike GANs, where images are generated from noise, a Markov chain learns a reverse diffusion process where Gaussian Noise is gradually added. For example:

x0 -> x1 -> x2 -> … -> xT

where …

x0 = the original image

x1 = x0 + Gaussian Noise

x2 = x1 + Gaussian noise, and so on.

Then a Neural Network is trained to generate the noise (reverse mean), and the process is reversed as follows to generate an image.

xT -> … -> x2 -> x1 -> x0

Vision Language Action Model: The VLA (Vision Language Action) model merges all the concepts of generative AI. For example, if we have an image I and some text T, the VLM (Vision Language Model) generates a text output (Y) by interpreting I and T. On the other hand, VLA models take I and T as input and generate a policy. You can write in the app, “Pick the black hammer from the table and give it to me”, and the VLA generates the policy to do so. We already know what a policy is from our previous lesson on Reinforcement Learning. Reinforcement Learning focuses on the learning mechanism, while VLA mainly focuses on the models to solve the problem. 

A famous VLA model from Google is called RT-2. RT-2 is an architecture in which pre-trained VLMs are co-fine-tuned using robotics and web datasets to generate instructions for robots. So far, we have talked about many models, but what are their use cases? This is one of the most important things an AI Generalist must know. Let me give you an example from the automotive world of how these generative models are used. Nowadays, automobiles are becoming increasingly autonomous. The Neural Networks running in the cars interpret pedestrians, signboards, etc., and take control actions based on that. How are these AI models trained? Definitely with real-world data. Say, for example, the AI model deployed in the car can read a signboard in a foggy scenario. From field data, we can have some clear images and some foggy winter images. With this, one can train the AI model. However, can we claim our models are reliable and robust? Definitely not. We would have more confidence if we could test our models with exactly 10%, 20%, 30%… 78% foggy images. But can we get field data? No, because we can’t generate natural fog. In this case, can we use Generative AI to generate the images and test our autonomous driving models? We can now, and this is one of the most important use cases of Generative AI in the automotive industry. Tools like CARLA use AI to generate scenarios.

So far, we have learned about complex models. In the next lessons, I will cover computational methods that help run AI models with less power.

khasnabish Avatar

Posted by

Leave a Reply

Discover more from AiSolve India

Subscribe now to keep reading and get access to the full archive.

Continue reading