Q1: What is Python? Why is it widely used in AI?
Python is a programming language used to give instructions to a computer.
You can think of it as a language that helps humans communicate with machines in a clear and readable way. It is called a high-level language because its syntax is close to human language and easy to understand.
Example:
print("Hello AI")
- It has powerful libraries like NumPy, Pandas, PyTorch, TensorFlow.
- It allows fast development and experimentation.
- It has excellent support for data handling.
- It integrates easily with APIs and external tools.
- In AI engineering, Python is used for:
- Data preprocessing
- Model training
- Model deployment
- API integration
- Automation
Interviewers expect you to understand that Python is not just easy — it is powerful enough for production systems.
Q2: What are Python’s built-in data types?
A data type defines what kind of value a variable stores.
For example:
Main built-in data types:
- int (Integer) Stores whole numbers.
age = 25
float
Stores decimal numbers.
price = 19.99
Floats are extremely important in AI because model weights and calculations use decimal values.
str (String)
Stores text.
name = "Chandra"
Strings are used in:
- NLP
- Prompts
- Chat systems
- API communication
- bool (Boolean) Stores True or False.
is_logged_in = True
Used in conditions and logic control.
list
Ordered collection of items.
numbers = [1, 2, 3]
- Dataset values
- Predictions
- Token lists
- tuple Ordered but cannot be modified after creation.
coordinates = (10, 20)
Useful when data should remain constant.
set
Unordered collection with no duplicates.
unique_ids = {1, 2, 3}
Useful for removing duplicate values.
dict (Dictionary)
Stores key-value pairs.
student = {
"name": "Chandra",
"score": 95
}
Dictionaries are heavily used in AI APIs and JSON responses.
Q3: What is dynamic typing?
Python is dynamically typed, meaning you do not need to declare the data type of a variable.
Example:
x = 10
x = "AI"
The same variable can hold different types at different times.
Python decides the type during runtime.
This flexibility makes Python very convenient for:
- Rapid development
- Data processing
- Handling different input types in AI systems
Q4: What is the difference between int and float?
An int stores whole numbers.
a = 5
A float stores decimal numbers.
b = 5.0
You can check the type:
print(type(5)) # int
print(type(5.0)) # float
- Neural network weights Probabilities Loss calculations
Q5: What is the difference between List and Tuple?
Both store multiple values, but they differ in mutability.
A list is mutable. That means it can be changed after creation.
my_list = [1, 2, 3]
my_list.append(4)
Now the list becomes:
[1, 2, 3, 4]
A tuple is immutable. That means it cannot be changed after creation.
my_tuple = (1, 2, 3)
Trying to modify it will result in an error.
Lists are used when data may change.
Tuples are used when data should remain fixed.
In AI, configuration values are sometimes stored as tuples because they should not change accidentally.
Q6: What is a Dictionary?
A dictionary stores data in key-value format.
It works like a real dictionary:
- You look up a word (key) to get its meaning (value).
Example:
student = {
"name": "Chandra",
"score": 95
}
Accessing value:
print(student["name"])
- JSON responses are dictionaries. Model outputs are often dictionaries. API responses are dictionaries. Understanding dictionaries is essential for AI engineers.
Q7: What is a Set and when should we use it?
A set is an unordered collection that does not allow duplicate values.
Example:
words = {"ai", "ml", "ai"}
print(words)
Output:
{"ai", "ml"}
Sets are useful when:
- Removing duplicates Tracking unique values Cleaning datasets in NLP
Q8: What is Mutability?
Mutability refers to whether an object can change after it is created.
Mutable objects:
- list dict set Example:
a = [1, 2]
b = a
a.append(3)
print(b)
Output:
[1, 2, 3]
Because both variables point to the same object in memory.
Immutable objects:
- int float str tuple Example:
x = 10
y = x
x = 20
Here, y is still 10 because integers are immutable.
Understanding mutability helps prevent unintended side effects in AI systems.
Q9: What is Type Casting?
Type casting means converting one data type to another.
Example:
x = "10"
y = int(x)
Now y is an integer.
Common conversions:
int()
float()
str()
list()
tuple()
- Converting user input Parsing API responses Preparing data before feeding it to models
Q10: What is the difference between == and is?
The operator == compares values.
a = [1, 2]
b = [1, 2]
print(a == b) # True
Because both lists contain the same values.
The operator is compares memory locations.
print(a is b) # False
Because they are two different objects in memory.
This distinction is often tested in interviews because many developers confuse value comparison with identity comparison.
No material available!
The trainer has not added any training content or tests to this lesson yet. Once the trainer adds content or tests, they will be displayed h
Q11: What is a function in Python? Why do we use it?
A function is a reusable block of code that performs a specific task.
Instead of writing the same code again and again, we define it once and reuse it.
Example:
def greet(name):
return "Hello " + name
Calling the function:
print(greet("Chandra"))
Output:
- Hello Chandra
- Why functions are important in AI:
- Breaking large systems into small reusable parts Creating model pipelines Writing reusable preprocessing logic Building modular AI systems
Q12: What is the difference between parameters and arguments?
Parameters are variables defined in the function definition.
Arguments are actual values passed when calling the function.
Example:
def add(a, b): # a and b are parameters
return a + b
add(2, 3) # 2 and 3 are arguments
This distinction is commonly tested in interviews.
Q13: What is the difference between return and print?
print() displays output on the screen.
return sends the result back to the caller.
Example:
def add(a, b):
print(a + b)
This only prints.
Now:
def add(a, b):
return a + b
This allows:
result = add(2, 3)
print(result)
- You must use return because functions are chained together in pipelines.
Q14: What are default arguments in Python?
Default arguments allow a function to work even if no value is passed.
Example:
def greet(name="User"):
return "Hello " + name
Calling:
print(greet())
Output:
- Hello User
- Used in AI:
- When building flexible model functions with optional parameters.
Q15: What is variable scope?
Scope defines where a variable is accessible.
There are mainly two types:
- Local Scope Defined inside a function.
def test():
x = 5
x cannot be accessed outside.
Global Scope
Defined outside functions.
x = 10
Accessible everywhere.
Understanding scope prevents bugs in AI pipelines.
Q16: What is the global keyword?
If you want to modify a global variable inside a function, you must use global.
Example:
x = 10
def change():
global x
x = 20
Without global, Python creates a new local variable.
Interview trap: Many beginners forget this and create bugs.
Q17: What are *args and **kwargs?
They allow flexible number of inputs.
args (multiple positional arguments)
def add(*args):
return sum(args)
print(add(1, 2, 3))
*kwargs (multiple keyword arguments)
def print_details(**kwargs):
return kwargs
print(print_details(name="AI", level="Beginner"))
- When writing flexible APIs or model wrappers.
Q18: What is a lambda function?
A lambda function is a small anonymous function written in one line.
Example:
square = lambda x: x * x
print(square(4))
Output:
16
Equivalent to:
def square(x):
return x * x
Used in:
- Sorting Data transformations Small utility functions
Q19: What is the difference between mutable default argument and safe default argument?
This is a very common interview question.
Problem example:
def add_item(item, my_list=[]):
my_list.append(item)
return my_list
Calling multiple times:
print(add_item(1))
print(add_item(2))
Output:
[1]
[1, 2]
Because the default list is shared across calls.
Safe version:
def add_item(item, my_list=None):
if my_list is None:
my_list = []
my_list.append(item)
return my_list
This avoids unexpected behavior.
Q20: What are some commonly used list methods?
Lists come with built-in methods that help manipulate data.
Example list:
numbers = [1, 2, 3]
Common methods:
append() → adds item at end
numbers.append(4)
numbers.insert(1, 10)
numbers.remove(2)
numbers.pop()
numbers.sort()
- Lists are used for storing predictions, tokens, batches of data.
Q21: What is the difference between append() and extend()?
append() adds entire object as one element.
a = [1, 2]
a.append([3, 4])
print(a)
Output:
[1, 2, [3, 4]]
extend() adds elements individually.
a = [1, 2]
a.extend([3, 4])
print(a)
Output:
[1, 2, 3, 4]
Q22: What is enumerate()?
enumerate() gives both index and value while iterating.
Example:
names = ["AI", "ML", "DL"]
for index, value in enumerate(names):
print(index, value)
Output:
- 0 AI 1 ML 2 DL
- Very useful in data preprocessing tasks.
Q23: What are dictionary methods?
Example dictionary:
student = {"name": "Chandra", "score": 95}
Common methods:
keys()
student.keys()
values()
student.values()
items()
student.items()
Used for looping:
for key, value in student.items():
print(key, value)
- We often loop through API responses which are dictionaries.
Q24: What is get() method in dictionary?
get() safely retrieves a value.
student.get("name")
If key does not exist:
student.get("age", "Not Found")
It avoids error.
Without get():
student["age"] # ❌ KeyError
Very important in AI when parsing unpredictable API responses.
Q25: What is a List Comprehension?
Short way to create lists.
Normal way:
squares = []
for i in range(5):
squares.append(i*i)
List comprehension:
squares = [i*i for i in range(5)]
More readable and efficient.
- Used for quick data transformations.
Q26: What is the difference between iterable and iterator?
Iterable:
- An object that can be looped over.
Examples:
- list tuple string dictionary Iterator:
- An object with __next__() method that produces next value.
Example:
numbers = [1,2,3]
it = iter(numbers)
This concept becomes important when learning:
- Generators Streaming large data Batch processing in AI
Q27: What is Object-Oriented Programming (OOP)?
OOP is a programming style where we organize code using classes and objects.
Instead of writing everything as separate functions, we group related data and behavior together.
Think of it like this:
- If you are building an AI Model system:
- Model name Model version Prediction function Training function All these belong together.
- So we create a class.
Q28: What is a Class?
A class groups two things in one place:
Data (attributes / properties)
→ what the object has
Examples: name, age, balance, color
Behavior (methods / functions inside a class)
→ what the object can do
Examples: deposit(), withdraw(), start_engine(), apply_discount()
So, a class is a way to model real-world entities in code.
Q29: Why Do We Need Classes?
Without classes, we usually store data in separate variables:
name1 = "Ravi"
age1 = 22
name2 = "Meena"
age2 = 24
This becomes messy when you have many records and behaviors.
With a class, you create a proper structure:
- One blueprint: Student
Q30: What is the init method?
__init__ is a special method (constructor).
It runs automatically when an object is created.
Example:
class Model:
def __init__(self, name):
self.name = name
model1 = Model("Resume Screener")
print(model1.name)
Output:
- Resume Screener
- __init__ is used to initialize object data.
- In AI:
- Store model name Store configuration Store API keys Store model parameters
Q31: What is self in Python?
self refers to the current object.
When we write:
class Model:
def __init__(self, name):
self.name = name
self.name means:
- Store this value inside this object.
If you create two objects:
model1 = Model("A")
model2 = Model("B")
Each object has its own name.
Without self, Python cannot differentiate between objects.
This is one of the most common beginner confusions in interviews.
Q32: What are Instance Variables and Methods?
Instance Variables
Variables that belong to an object.
Example:
class Student:
def __init__(self, name):
self.name = name
Here:
- name is instance variable.
- Instance Methods Functions inside a class.
Example:
class Model:
def __init__(self, name):
self.name = name
def predict(self):
return "Prediction from " + self.name
Calling:
model = Model("AI Model")
print(model.predict())
predict()
train()
evaluate()
Q33: What is Inheritance?
Inheritance allows one class to reuse another class.
Example:
class BaseModel:
def train(self):
return "Training"
class AIModel(BaseModel):
def predict(self):
return "Predicting"
Here:
- AIModel inherits from BaseModel.
Now:
model = AIModel()
print(model.train())
Output:
- Training
- Used in AI:
- Base model class Specialized models Custom retrievers Custom agents Inheritance reduces code duplication.
Q34: What is Polymorphism?
Polymorphism means:
- Same method name behaves differently for different objects.
Example:
class Cat:
def sound(self):
return "Meow"
class Dog:
def sound(self):
return "Bark"
Both have sound() method, but behavior is different.
- Different models may have same method:
model.predict()
But internal logic differs.
Q35: What is Encapsulation?
Encapsulation means hiding internal details and exposing only what is necessary.
Example:
class BankAccount:
def __init__(self, balance):
self.__balance = balance
def get_balance(self):
return self.__balance
__balance is private.
You cannot access:
- account.__balance ❌
- This protects internal data.
- In AI:
- Protect model internals Protect API keys Protect sensitive configuration
Q36: What are Dunder (Magic) Methods?
Dunder means "double underscore".
Examples:
- __init__ __str__ __len__ __repr__ Example:
class Model:
def __init__(self, name):
self.name = name
def __str__(self):
return f"Model: {self.name}"
model = Model("AI")
print(model)
Output:
Model: AI
These methods allow you to customize object behavior.
- Custom logging Debug printing Object comparison Overloading operators
Q37: What are the different file modes in Python?
When opening a file, we specify a mode:
- Mode Meaning "r" Read "w" Write (overwrites file) "a" Append "r+" Read + Write "b" Binary mode Example:
file = open("data.txt", "w")
Interview tip:
- Always mention that "w" overwrites existing content.
Q38: What is the recommended way to open files?
We use the with statement.
Example:
with open("example.txt", "r") as file:
content = file.read()
Why is this better?
Because:
- It automatically closes the file. Prevents memory leaks. Cleaner and safer. In AI systems:
- Always use with for file handling.
Q39: How do you read a file line by line?
Example:
with open("example.txt", "r") as file:
for line in file:
print(line.strip())
Useful when:
- Reading large datasets Processing logs Processing training data Reading line by line saves memory.
Q40: How do you write data to a file?
Example:
with open("output.txt", "w") as file:
file.write("Hello AI")
Append mode:
with open("output.txt", "a") as file:
file.write("\\nNew Line")
- Saving predictions Writing experiment results Logging outputs
Q41: How do you work with JSON files?
- API responses are JSON Configuration files are JSON Model outputs are JSON Import JSON module:
import json
Writing JSON:
data = {"name": "AI", "score": 95}
with open("data.json", "w") as file:
json.dump(data, file)
Reading JSON:
with open("data.json", "r") as file:
data = json.load(file)
print(data["name"])
Interview tip:
Always differentiate between:
- json.dump() → write to file json.dumps() → convert to string
Q42: What is a Module in Python?
A module is a Python file containing functions, classes, or variables.
Example:
- Create file: math_utils.py
def add(a, b):
return a + b
Import it:
import math_utils
print(math_utils.add(2, 3))
Modules help in:
- Organizing large AI projects Reusing code Maintaining clean structure
Q43: What are the different import styles?
There are multiple ways to import.
Basic import
import math
Import specific function
from math import sqrt
Import with alias
import numpy as np
Aliasing is common in AI:
numpy as np pandas as pd torch as torch Interviewers expect you to know these styles.
Q44: What is name == "main" ?
Every Python file has a built-in variable called __name__.
If you run a file directly:
print(__name__)
Output:
- __main__
If you import it:
- __name__ becomes the module name.
Example:
if __name__ == "__main__":
print("This runs only when file is executed directly")
This prevents certain code from running during import.
- Useful when writing scripts that can be both:
- Imported as module Run as standalone program
Q45: What is an Exception in Python?
An exception is an error that occurs during program execution.
There are two types of errors:
- Syntax Error Happens when code is written incorrectly.
Example:
- print("Hello"
- This will not even run.
- Exception (Runtime Error) Happens while the program is running.
Example:
x = 10 / 0
This causes:
- ZeroDivisionError
- In AI systems, runtime errors are very common because:
- User input may be wrong API may fail File may be missing That is why exception handling is critical.
Q46: What is try-except block?
The try-except block allows you to handle errors gracefully.
Basic syntax:
try:
x = 10 / 0
except ZeroDivisionError:
print("Cannot divide by zero")
Instead of crashing, the program continues safely.
This prevents your application from crashing.
Interview tip:
- Always mention that exception handling improves robustness and reliability.
Q47: What is the difference between except and finally?
except
Runs only if an error occurs.
finally
Runs no matter what.
Example:
try:
file = open("data.txt", "r")
except FileNotFoundError:
print("File not found")
finally:
print("Execution finished")
finally is useful for:
- Closing files Releasing resources Cleaning up connections In AI:
- You may close database connections or API sessions in finally.
Q48: What is raise in Python?
The raise keyword is used to manually trigger an exception.
Example:
def check_age(age):
if age < 0:
raise ValueError("Age cannot be negative")
Why use raise?
Because:
- You want to enforce business rules You want to stop invalid data from entering system In AI:
You may raise error if:
Input format is invalid Required field is missing Model configuration is incorrect Interviewers like candidates who understand defensive programming.
Q49: What are Custom Exceptions?
You can create your own exception class.
Example:
class InvalidModelError(Exception):
pass
raise InvalidModelError("Model not supported")
Why useful?
In large AI systems:
- You want meaningful errors You want to differentiate between different failure types You want cleaner debugging Example:
try:
load_model("abc")
except InvalidModelError as e:
print("Custom error:", e)
This makes production systems easier to debug.
Q50: What is the Global Interpreter Lock (GIL)? How does it impact multi-threading vs multi-processing during heavy data preprocessing?
The Global Interpreter Lock (GIL) is a mechanism in CPython that ensures only one thread executes Python bytecode at a time. This prevents multiple threads from executing Python code simultaneously on multiple CPU cores.
In AI engineering, heavy data preprocessing (e.g., tokenizing text, resizing images, or scaling arrays) is highly CPU-bound. Multi-threading in Python will NOT speed up these tasks because of the GIL. Instead, threads will fight for lock acquisition, making it even slower. To utilize all CPU cores for preprocessing, you must use multi-processing (which runs separate Python processes with their own GILs) or offload calculations to compiled C/C++ backends (like NumPy or PyTorch) which release the GIL.
Example:
import multiprocessing
import numpy as np
def preprocess_chunk(chunk):
# Perform heavy matrix calculations that run in parallel
return np.sin(chunk) * np.cos(chunk)
if __name__ == "__main__":
data = np.random.rand(1000000)
chunks = np.array_split(data, 4)
with multiprocessing.Pool(processes=4) as pool:
results = pool.map(preprocess_chunk, chunks)
Interview tip: Always mention that the GIL is a CPython implementation detail. Explain that multi-threading is great for I/O-bound tasks (like scraping data or calling API endpoints), but multi-processing is mandatory for CPU-bound tasks in Python.
Q51: What is the difference between a shallow copy and a deep copy? When is deep copy critical for model configurations and weight vectors?
A shallow copy creates a new collection object, but inserts references to the original objects. If you modify a nested object inside a shallow copy, the change will reflect in the original.
A deep copy creates a new collection object and recursively copies all nested objects inside it, ensuring complete isolation.
Example:
import copy
original_config = {"model": "transformer", "params": {"layers": 12, "lr": 0.001}}
# Shallow Copy
shallow_config = copy.copy(original_config)
shallow_config["params"]["lr"] = 0.01 # Modifies original_config!
# Deep Copy
deep_config = copy.deepcopy(original_config)
deep_config["params"]["lr"] = 0.0001 # Safe, original_config is untouched.
When modifying configurations, fine-tuning hyperparameters, or cloning model weights in a reinforcement learning loop (e.g. copying target networks in DQN), you must use deep copy. A shallow copy will cause weights or params to be modified globally, ruining the training loop.
Interview tip: Explain that mutable nested objects (like dictionaries inside lists or nested configurations) are not duplicated by default in a shallow copy. If the interviewer asks about tensors, note that PyTorch uses `.clone().detach()` for safe copying of weights instead of Python's copy library.
Q52: What is the difference between a list comprehension and a generator expression in terms of memory utilization? How does this impact processing 10M+ tokens?
A list comprehension computes all elements immediately and stores the entire list in memory.
A generator expression calculates each element on demand (lazy evaluation), utilizing almost zero memory at initialization.
Example:
# List comprehension: loads all squares into memory
squares_list = [x**2 for x in range(1000000)]
# Generator expression: yields elements one by one
squares_gen = (x**2 for x in range(1000000))
When tokenizing massive datasets (like a 10M+ token corpus), loading the entire processed output into a list comprehension will crash the container due to Out-Of-Memory (OOM) errors. Using a generator allows you to stream tokens line-by-line, passing them to the training loop dynamically without building the entire array in RAM.
Interview tip: Demonstrate your understanding by mentioning the `sys.getsizeof()` function to show the memory difference. A list of 1 million items takes megabytes of memory, whereas the generator object takes only a few bytes.
Q53: How do you create custom Generators in Python to build memory-efficient streaming dataloaders for infinite datasets?
You create a custom generator using a function with the yield keyword. Unlike return, yield pauses the function execution, returns the value, and resumes from the exact same state on the next request.
Example:
import time
def stream_batches(dataset_path, batch_size=32):
with open(dataset_path, "r") as file:
batch = []
for line in file:
batch.append(line.strip())
if len(batch) == batch_size:
yield batch
batch = []
if batch:
yield batch
# Streaming items
for batch in stream_batches("large_corpus.txt", batch_size=2):
print("Feeding batch to model:", batch)
Generators are the foundation of dataloaders in PyTorch and TensorFlow. When training models on datasets that are larger than the available RAM (e.g., hundreds of gigabytes of text or images), a generator streams data from disk dynamically in batches, keeping memory usage constant.
Interview tip: Explain that generators implement the Iterator Protocol (they have `__iter__` and `__next__` methods automatically), which allows them to be used directly in loops.
Q54: What is the difference between CPU-bound and I/O-bound tasks in Python? Which concurrency approach (asyncio, threading, or multiprocessing) should be used for model inference vs API data fetching?
CPU-bound tasks spend most of their time doing calculations on the CPU. Examples: Matrix multiplications, image resizing, tokenization.
I/O-bound tasks spend most of their time waiting for external operations. Examples: Querying databases, downloading datasets, calling LLM API endpoints.
Model inference is highly CPU/GPU bound. For local running models, you should use multiprocessing to bypass the GIL, or offload it to standard C++ backends (like PyTorch C++ engines).
API data fetching (like calling the OpenAI API for 1000 prompts) is I/O-bound. You should use asyncio or multi-threading because the program spends 99% of its time waiting for the network response.
Example of I/O Bound Concurrency:
import asyncio
import aiohttp
async def fetch_prediction(session, prompt):
async with session.post("https://api.openai.com/v1/completions", json={"prompt": prompt}) as response:
return await response.json()
async def main():
async with aiohttp.ClientSession() as session:
tasks = [fetch_prediction(session, f"Prompt {i}") for i in range(10)]
results = await asyncio.gather(*tasks)
Interview tip: Explain that for CPU-bound tasks, multiprocessing creates independent OS processes, which incurs overhead. For I/O-bound tasks, asyncio utilizes a single thread with an event loop, making it lightweight and highly scalable compared to spawning threads.
Q55: What are Python Decorators? How can we use decorators to create a execution-timer/logger for tracking model training and inference latency?
A decorator is a function that takes another function as an argument, extends its behavior without modifying it, and returns the modified function.
Example:
import time
import functools
def log_latency(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
start_time = time.perf_counter()
result = func(*args, **kwargs)
end_time = time.perf_counter()
print(f"Latency of {func.__name__}: {end_time - start_time:.4f} seconds")
return result
return wrapper
@log_latency
def run_inference(inputs):
time.sleep(0.5) # Simulate model inference
return "Prediction"
run_inference("test_input")
Decorators are heavily used in production AI systems to log inputs/outputs, monitor model prediction latency, verify API authentication, or catch connection retries without cluttering core model code.
Interview tip: Always mention `@functools.wraps(func)` inside your custom decorator. Without it, the decorated function will lose its original name and docstring (e.g. `run_inference.__name__` would print `wrapper`), which ruins debugging and logging in production.
Q56: What is a Context Manager (with statement) and how do you implement a custom one to safely manage GPU memory or file access?
A context manager is an object that defines the runtime context to be established when executing a with statement. It uses `__enter__` and `__exit__` methods to setup and teardown resources.
Example:
class GPUMemoryManager:
def __init__(self, device):
self.device = device
def __enter__(self):
print(f"Allocating GPU Memory on device {self.device}")
return self
def __exit__(self, exc_type, exc_val, exc_tb):
print(f"Releasing GPU Memory on device {self.device}")
# Clean up memory caches (e.g. torch.cuda.empty_cache())
if exc_type:
print(f"Exception occurred: {exc_val}")
return False # Do not suppress exceptions
with GPUMemoryManager("cuda:0") as manager:
print("Running matrix multiplication...")
# Matrix operations go here
Context managers are vital in AI pipelines to prevent GPU memory leaks. For example, PyTorch uses context managers like `with torch.no_grad()` to temporarily disable gradient calculations (saving massive GPU memory during inference) and `with torch.cuda.device(1)` to switch active cards safely.
Interview tip: Be prepared to write a context manager using the `@contextlib.contextmanager` decorator as a cleaner alternative. It uses generators instead of defining class methods.
Q57: How does Python's zip() and map() functions help optimize dataset alignment (e.g., matching inputs with target labels) without nested loops?
The `zip()` function joins elements from two or more iterables position-by-position, returning tuples.
The `map()` function applies a specific function to all items in an input iterable.
Example:
# Pairing data using zip()
inputs = ["Describe AI", "What is NLP?"]
labels = ["RAG", "Text Parsing"]
paired_data = list(zip(inputs, labels))
# Processing using map()
def clean_text(text):
return text.strip().lower()
clean_prompts = list(map(clean_text, [" PromptA ", "PromptB "]))
Using nested `for` loops in Python is extremely slow. `zip()` and `map()` provide memory-efficient iterations. `zip()` is used to bind input tokens with corresponding labels (like in Named Entity Recognition tasks), and `map()` is used to clean dataset lists instantly before tokenization.
Interview tip: Explain that both `zip()` and `map()` in Python 3 return iterators (lazy evaluation) rather than lists. They do not allocate memory for the final paired output until you iterate over them or cast them (e.g., via `list()`).
Q58: What are any() and all() built-ins? How do you use them to validate threshold filters or data sanity masks in dataset pipelines?
The `any()` function returns True if at least one element in an iterable is truthy.
The `all()` function returns True only if all elements in an iterable are truthy.
Example:
predictions_scores = [0.85, 0.92, 0.45, 0.76]
# Check if any prediction is below the confidence threshold
has_low_confidence = any(score < 0.50 for score in predictions_scores)
print("Low confidence warning:", has_low_confidence)
# Check if all predictions meet minimum quality checks
meets_quality = all(score > 0.40 for score in predictions_scores)
print("Passes pipeline check:", meets_quality)
In validation pipelines, `all()` is used to verify that no labels are missing (e.g., checking that all tokens are non-null). `any()` is useful for outlier detection, warning you if any single feature inside a multi-dimensional array falls outside acceptable boundaries.
Interview tip: Highlight that both functions implement short-circuiting: `any()` stops evaluation the moment it finds the first True value, and `all()` stops immediately at the first False value. This makes them highly optimized for large sequences.
Q59: Why are Python Type Hints (static typing declarations via typing) and toolsets like Pydantic essential when constructing configuration schemas for LLM Agents?
Python type hints let you declare the expected type of variables, function parameters, and return values. Pydantic leverages these type hints at runtime to validate data structures and raise validation errors if types mismatch.
Example:
from pydantic import BaseModel, Field
class AgentConfig(BaseModel):
agent_name: str
temperature: float = Field(default=0.7, ge=0.0, le=1.0)
max_tokens: int
# Valid configuration
config = AgentConfig(agent_name="RAGBot", temperature=0.5, max_tokens=150)
# Invalid configuration -> Raises ValidationError
try:
bad_config = AgentConfig(agent_name="FailBot", temperature=1.5, max_tokens="lots")
except Exception as e:
print(e)
LLM responses are unpredictable strings. When building AI agents, you need to parse unstructured LLM outputs into structured objects. Pydantic acts as the parser and guardrail: defining a Pydantic schema forces the LLM (via JSON mode or structured outputs) to return matching types, preventing parsing crashes in production.
Interview tip: Point out that standard Python type hints are purely for IDE analysis and static type checkers (like `mypy`) — they do not enforce types at runtime. Pydantic turns type hints into runtime validations.
Q60: How do you handle OpenAI/LLM API rate limits dynamically in Python using exponential backoff exceptions?
Exponential backoff is a standard error handling strategy where you wait progressively longer between retries of a failed network request to avoid overloading the API server.
Example:
import time
import random
def call_llm_with_backoff(prompt, max_retries=5):
base_delay = 1.0 # start with 1 second delay
for attempt in range(max_retries):
try:
# Simulate API call which might raise RateLimitError
if random.random() < 0.7:
raise Exception("Rate limit exceeded")
return "Successful prediction"
except Exception as e:
if attempt == max_retries - 1:
raise e
# Calculate backoff time with jitter (randomness) to avoid thundering herd
delay = (base_delay * (2 ** attempt)) + random.uniform(0, 1)
print(f"Attempt {attempt + 1} failed. Retrying in {delay:.2f}s...")
time.sleep(delay)
call_llm_with_backoff("Translate: Hello")
Production RAG systems call external APIs (OpenAI, Anthropic, Pinecone) in loops. Spawning dozens of agents simultaneously can instantly trigger Rate Limit exceptions. Implementing exponential backoff with jitter prevents crashes and ensures your system recovers automatically.
Interview tip: Mention library integrations like tenacity (`@retry(wait=wait_random_exponential(...))`). Interviewers love when candidates know how to use industry-standard production packages instead of writing raw delay loops from scratch.
Q61: What is the difference between Python's virtual environments (venv), conda, and modern package managers like poetry or uv when shipping AI dependencies across CPU vs GPU environments?
`venv` is Python's built-in tool that manages standard package isolates in a virtual environment using `pip` (limited to PyPI packages).
`conda` is a cross-platform package manager that can install Python and binary libraries (like CUDA, C++ compilers, and GPU drivers) directly.
`poetry` and `uv` are modern lockfile-based dependency managers that resolve dependencies deterministically, generating locked lockfiles to guarantee the exact same packages are installed on staging and production.
Deep learning environments require matching specific PyTorch versions with correct CUDA drivers. Spawning a container using just a basic pip `requirements.txt` often installs incompatible CPU versions of PyTorch. Conda is used to manage system-level binary dependencies (like CUDA), while poetry/uv is used in microservices to ensure fast, deterministic package replication.
Interview tip: Highlight that `uv` is written in Rust and is 10-100x faster than standard `pip` or `poetry`. Mention that lockfiles (`poetry.lock`, `uv.lock`) are critical in production to prevent dependencies from changing silently between deployments.