Using an AI Library Responsibly
Chapter Eighty-Eight
Syllabus topic Module 2, "Demo of OpenAI/TensorFlow Tools"
Pages 581 to 586 of 591
In one line
A library call is one of the algorithms in this book with the arguments hidden, and using it responsibly means knowing which algorithm, what it assumed, and what to check before believing the number it returns.
MU's paired practical, Computer Science Practical 5, sets a demonstration using OpenAI or TensorFlow tools, and the machine learning practicals use scikit-learn throughout. This chapter is about what those calls are doing and what must be checked. It is not a tutorial in any library's syntax, which changes, and which the library's own documentation gives better.
A declaration about this chapter
Every other program in this book was run on three interpreters and its output pasted from the run. The listings here are not. scikit-learn, TensorFlow and the OpenAI client are not installed on the machine this book was checked on, so nothing below could be executed, and no output is claimed for any of it.
Each block is therefore written as a form: the shape of the call, with the parts that matter named. Check the library's current documentation before using any of them, because these interfaces change between versions and a printed line in a book is out of date the moment a release happens.
Every call is a chapter of this book
This table is the chapter. A student who can complete it can read a practical's solution; one who cannot is copying.
| The call | What it actually is | Chapter |
|---|---|---|
LinearRegression | least squares by the normal equations | The First Learner: Fitting a Straight Line |
SGDRegressor, SGDClassifier | gradient descent, in batches | Gradient Descent |
Ridge, Lasso | least squares with a penalty on the weights | Regularisation |
DecisionTreeClassifier | ID3 or CART: choose the split with the best information gain | Building a Decision Tree with ID3 |
RandomForestClassifier | bagging plus a random subset of features per split | Ensemble Methods, Bagging and The Random Forest |
AdaBoostClassifier | the reweighting of Boosting and AdaBoost | |
GaussianNB, MultinomialNB | the independence assumption, with smoothing | Naive Bayes |
KNeighborsClassifier | store everything, compare at query time | k-Nearest Neighbours |
SVC(kernel=...) | the maximum margin, with the kernel trick | The Kernel Trick |
KMeans | assign, recentre, repeat | Clustering and k-means |
AgglomerativeClustering(linkage=...) | merge the two closest clusters | Hierarchical Clustering and Judging a Clustering |
train_test_split, cross_val_score | holding data back, and k folds | Evaluating a Model |
classification_report | precision, recall, F1, per class | Evaluating a Model |
StandardScaler | the scaling that distance-based methods need | k-Nearest Neighbours |
A Sequential model with Dense layers | a feed-forward network | Neural Networks and The Perceptron |
loss='categorical_crossentropy', optimizer='adam' | the loss and the descent rule | Backpropagation, Gradient Descent |
model.fit(..., epochs=..., validation_split=...) | training, with a held-out set to watch for overfitting | Overfitting and Underfitting |
Two entries deserve a second look.
Using an AI Library Responsibly
SVC(kernel='rbf') is a decision, not a default. The Kernel Trick showed what a kernel does and what it costs. Choosing one without being able to say what shape of boundary it allows is choosing at random.
RandomForestClassifier gives up interpretability. Transparency and Explainability measured the trade: one readable tree at 0.6657 against an unreadable forest at 0.7450. That is a real gain and it is a real loss, and the loss is invisible until someone asks why a person was refused.
The shape of a scikit-learn program
X, y = features, labels # X: rows by columns, y: one label per row
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
scaler.fit(Xtr); Xtr = scaler.transform(Xtr); Xte = scaler.transform(Xte)
model = SomeClassifier(**chosen_hyperparameters)
model.fit(Xtr, ytr) # the learning
print(classification_report(yte, model.predict(Xte)))
print(confusion_matrix(yte, model.predict(Xte)))Four things in that form carry the whole responsibility of the chapter, and every one of them is a line a beginner deletes.
stratify=y. Without it, an unbalanced problem can put almost none of the rare class in the test set, and Evaluating a Model showed what unbalanced data does to an accuracy figure.
random_state=0. Without a fixed seed the split changes every run, so the score changes every run, and a comparison between two models measures the split rather than the models.
scaler.fit(Xtr) and not scaler.fit(X). Fitting the scaler on all the data lets the test set's mean and spread into the training, which is data leakage: the score improves and the model does not. The rule is absolute: anything fitted must be fitted on the training data alone.
classification_report and confusion_matrix rather than score. Evaluating a Model measured a classifier at 0.9700 accuracy that was worthless. A single number is not a result.
The shape of a TensorFlow or Keras program
model = Sequential([Dense(32, activation='relu'), Dense(n_classes, activation='softmax')])
model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
history = model.fit(Xtr, ytr, epochs=50, validation_split=0.2, verbose=0)
# the two curves in history are the point: training loss falling while
# validation loss rises IS overfitting, seen as it happens
loss, acc = model.evaluate(Xte, yte)The comment is the content. The validation curve turning upward while the training curve keeps falling is the picture Overfitting and Underfitting drew, and it is the one thing a practical write-up should contain and usually does not.
And the honest note about epochs=50: it is a guess. Early stopping, halting when the validation loss stops improving, is the principled version, and it is one argument to the same call.
The shape of a call to a hosted model
key = os.environ["OPENAI_API_KEY"] # NEVER written in the source file
reply = client.responses.create(model=..., input=prompt)
# everything in `prompt` has left the machine. treat it as published.Using an AI Library Responsibly
Three rules, and the first two are absolute.
The key is a secret. It goes in an environment variable or a secrets manager, never in the source and never in a repository. A key in a public repository is found by automated scanners in minutes and billed to its owner.
The prompt leaves the machine. Anything sent is disclosed to a third party: a classmate's marks, a patient record, a friend's message, an unpublished manuscript. Treat the prompt as published, because for the purposes of the person whose data it was, it is.
The reply is unverified. Hallucination in Generative AI measured why: the model scores how ordinary a sequence of words looks and a false sentence of ordinary phrases scores well. Every specific claim, number, citation and date in a reply must be checked against a source before it is used.
The checks before a result is believed
Ten, in order, and this list is the answer to "how would you validate your practical's result".
| Check | Why | |
|---|---|---|
| 1 | What is the baseline? The largest class's share, or the simplest rule | Evaluating a Model: 0.9700 was the floor, not the result |
| 2 | Is the test data genuinely held out, and was nothing fitted on it? | data leakage inflates every number |
| 3 | Are there duplicate rows across the split? | the same row in train and test is memorised, not learned |
| 4 | Does any feature encode the answer? | a column recorded after the outcome is a leak |
| 5 | What is the confusion matrix, and which error costs more? | the two errors are not the same error |
| 6 | What are the scores per group? | Bias and Fairness in AI Models |
| 7 | What is the spread across folds, not just the mean? | Evaluating a Model: five folds ran 0.50 to 1.00 |
| 8 | Were the hyperparameters chosen on the test set? | then the test set is training data and the score is meaningless |
| 9 | Does the result survive a different random seed? | if not, it is the seed's result |
| 10 | Is the training score being reported by mistake? | it means nothing; the model has seen those rows |
Check 8 is the one that quietly ruins student projects. Trying twenty models and reporting the best test score reports the maximum of twenty noisy numbers. Hyperparameters are chosen on a validation set or by cross validation inside the training data, and the test set is touched once, at the end.
Responsibilities that are not about accuracy
| Licence | a model and a data set each have one, and "available to download" is not a licence to use |
| Provenance of the data | who is in it, and were they asked |
| Personal data | India's Digital Personal Data Protection Act, 2023 applies to a student project as much as to a company |
| Attribution | say which library, which version, and which pretrained weights |
| Reproducibility | record the versions and the seed, or the result is not a result |
| Cost and footprint | Ethical Issues in AI Systems: one large training run was estimated at 284 tonnes of CO2 |
| Scope | a model trained on one population is not evidence about another |
Using an AI Library Responsibly
Distinctions
| Calling a library | Understanding the call | |
|---|---|---|
| Needs | the documentation | the chapter |
| Produces | a number | a number you can defend |
| Examinable | rarely | yes |
fit on the training data | fit on everything | |
|---|---|---|
| Score | honest | inflated |
| Name for the error | data leakage |
score | classification_report | |
|---|---|---|
| Returns | one number | precision, recall, F1, per class |
| Safe on unbalanced data | no | yes |
What it does not mean
A library call is not a black box. Every one in the table is an algorithm in this book.
A default is not a choice. A kernel, a depth, a learning rate and a number of epochs are all decisions.
A high score is not a result. Compare it with the baseline first.
Fitting a scaler on all the data is not harmless. It is leakage and it inflates every number afterwards.
Choosing a model on the test set is not evaluation. It makes the test set training data.
An API key in a file is not private. It is found by scanners within minutes of being published.
A prompt is not private either. Treat it as published.
A model's reply is not a source. Check every specific claim against one.
Quick revision
- Every library call is a chapter of this book.
LinearRegression,SGDClassifier,Ridge,DecisionTreeClassifier,RandomForestClassifier,AdaBoostClassifier,GaussianNB,KNeighborsClassifier,SVC,KMeans,AgglomerativeClustering,cross_val_score,StandardScaler,SequentialwithDense. - The four lines that carry the responsibility:
stratify,random_state, fit the scaler on the training data only, and report a confusion matrix, not a score. - Data leakage: anything fitted must be fitted on the training data alone.
- In Keras, the training and validation curves are the result: validation rising while training falls is overfitting. Early stopping replaces a guessed epoch count.
- Hosted models: the key is a secret, the prompt leaves the machine and should be treated as published, and the reply is unverified.
- The ten checks: baseline, genuinely held-out test data, duplicates, a feature encoding the answer, the confusion matrix, per-group scores, the spread across folds, hyperparameters not chosen on the test set, survival of a different seed, and not reporting the training score.
- Choosing among twenty models by test score reports the maximum of twenty noisy numbers. Tune on a validation set; touch the test set once.
- Beyond accuracy: licence, data provenance, personal data under the DPDP Act, 2023, attribution, reproducibility (versions and seed), cost, and scope.
Using an AI Library Responsibly
Test yourself
1. A practical uses RandomForestClassifier. What is it doing, in the terms of this book? Growing many decision trees, each on a bootstrap sample of the rows drawn with replacement and each restricted at every split to a random subset of the features, then combining them by majority vote. The random feature subset makes the trees less alike so that their errors cancel, at the cost of each tree being individually worse and of the whole model no longer being readable as a set of rules.
2. Why must a scaler be fitted on the training data only? Because fitting it on all the data uses the test set's mean and spread to transform the training data, so information from the test set enters the model. That is data leakage: the reported score improves while the model does not, and the improvement disappears on genuinely new data.
3. Name four things in a scikit-learn pipeline a beginner omits and say why each matters. Stratifying the split, without which an unbalanced problem may place almost none of the rare class in the test set; fixing the random seed, without which the score changes every run and comparisons measure the split rather than the model; fitting transformations on the training data alone, without which there is leakage; and reporting a confusion matrix and per-class report rather than a single accuracy, since an accuracy of 0.97 was shown earlier to belong to a worthless classifier.
4. What should a practical write-up show from a Keras training run? The training and validation curves together. Training loss falling while validation loss rises is overfitting, visible as it happens, and it is the one piece of evidence that shows whether the chosen number of epochs was right. Early stopping on the validation loss replaces the guess.
5. Give three rules for calling a hosted model. The API key is a secret and belongs in an environment variable or secrets manager, never in source or a repository, since published keys are found by automated scanners within minutes. Everything in the prompt leaves the machine and should be treated as published, so another person's data must not be sent without their consent. And the reply is unverified text produced by a process that scores how ordinary a word sequence looks, so every specific claim, number, citation and date must be checked against a source.
Using an AI Library Responsibly
6. Why is reporting the best test score among twenty tried models misleading? Because it reports the maximum of twenty noisy measurements, which is biased upward, and because each comparison used the test set to make a choice, so the test set has become part of the training process. Hyperparameters must be selected on a validation set or by cross validation within the training data, and the test set used once, at the end.
7. List the checks you would make before believing a classification result. Compare it with the baseline given by the largest class or the simplest rule; confirm the test data was genuinely held out and that nothing was fitted on it; look for duplicate rows across the split; check that no feature records the answer or something recorded after it; read the confusion matrix and decide which error costs more; compute the scores separately for each group; report the spread across folds and not only the mean; confirm the hyperparameters were not chosen on the test set; repeat with a different seed; and make sure the number quoted is not the training score.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.