If you use tf.keras.utils.to_categorical to one-hot the label vector, the integers should start from 0 to num_classes, source. In your case, you should do as follows

import tensorflow as tf 
import numpy as np 

a = np.array([1,2,4,3,5,2,4,2,1])
y_tf = tf.keras.utils.to_categorical(a-1, num_classes = 5)
y_tf

array([[1., 0., 0., 0., 0.],
       [0., 1., 0., 0., 0.],
       [0., 0., 0., 1., 0.],
       [0., 0., 1., 0., 0.],
       [0., 0., 0., 0., 1.],
       [0., 1., 0., 0., 0.],
       [0., 0., 0., 1., 0.],
       [0., 1., 0., 0., 0.],
       [1., 0., 0., 0., 0.]], dtype=float32)

or, you can use pd.get_dummies,

import pandas as pd 
import numpy as np 

a = np.array([1,2,4,3,5,2,4,2,1])
a_pd = pd.get_dummies(a).astype('float32').values 
a_pd

array([[1., 0., 0., 0., 0.],
       [0., 1., 0., 0., 0.],
       [0., 0., 0., 1., 0.],
       [0., 0., 1., 0., 0.],
       [0., 0., 0., 0., 1.],
       [0., 1., 0., 0., 0.],
       [0., 0., 0., 1., 0.],
       [0., 1., 0., 0., 0.],
       [1., 0., 0., 0., 0.]], dtype=float32)
Answer from Innat on Stack Overflow
🌐
Keras
keras.io › api › layers › preprocessing_layers › categorical › category_encoding
Keras documentation: CategoryEncoding layer
Values can be "one_hot", "multi_hot" or "count", configuring the layer as follows: - "one_hot": Encodes each individual element in the input into an array of num_tokens size, containing a 1 at the element index.
Discussions

classification - How to use one hot encoding of string categorical features in keras? - Data Science Stack Exchange
I am dealing with a binary classification problem. The output column of my dataset is already encoded in 0/1. The problem is that I have many categorical features (columns), which are strings and I... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
neural network - Beyond one-hot encoding for LSTM model in Keras - Data Science Stack Exchange
I have an LSTM model in Keras for categorical classification (20 possible categories). In many cases, my data can fit multiple categories. Obviously, my current model uses one-hot encoding and fi... More on datascience.stackexchange.com
🌐 datascience.stackexchange.com
How can I decode one hot vector? (and when can we use it?)
np.argmax(one_hot, axis=1) More on reddit.com
🌐 r/MachineLearning
2
0
September 12, 2016
What are embedding layers?
Words are essentially categorical variables, with the total number of classes being your vocab size, and each word being a category within it. Representing words as integers by assigning a number to each category doesn't really make sense as integers have magnitude, while one word can't really be "larger" than another. Using the traditional categorical representation of one-hot vectors is impractical because for any reasonably sized vocab size (~10000), it's going to be extremely sparse, and would present significant computational issues as well, especially for deep NNs. So the solution is to map each variable to a fixed-size vector that's much lower dimension and much less sparse. These mapped vectors can then be fed into your model. This is the main advantage for using emebdding layers (like in Keras), it's a trained mapping that reduces the dimensionality of categorical variables. Word embeddings in general (including word2vec) just produce some kind of mapping from words to vectors. Keras embedding layers do this just as a simple mapping that's trained, like looking up a word in a dictionary. Word2vec trains a neural network that produces this mapping. There are many advantages to this, including a semantic representation of words as you described. The advantage for NNs, in addition to dimensionality reduction can also be thought of as a sort of feature extractor, using the underlying meaning of words as input instead of just variables, especially if you use pretrained embeddings. More on reddit.com
🌐 r/learnmachinelearning
13
4
August 1, 2020
Top answer
1 of 2
3

For string data, use get_dummies() (from Pandas). to_categorical() takes integers as inputs.

There are two important differences between Keras: to_categorical() and Pandas: get_dummies().


Keras: to_categorical()

  • to_categorical() takes integers as input (no strings allowed).
  • to_categorical() generates dummies starting at 0 by default!

Looking at the help function:

print(help(to_categorical))

Says:

to_categorical(y, num_classes=None, dtype='float32')
    Converts a class vector (integers) to binary class matrix.

    E.g. for use with categorical_crossentropy.

    # Arguments
        y: class vector to be converted into a matrix
            (integers from 0 to num_classes).
        num_classes: total number of classes.
        dtype: The data type expected by the input, as a string
            (`float32`, `float64`, `int32`...)
...

So if your data is numeric (int), you can use to_categorical(). You can check if your data is an np.array by looking at .dtype and/or type().

import numpy as np
npa = np.array([2,2,3,3,4,4])
print(npa.dtype, type(npa))
print(npa)

Result:

int32 <class 'numpy.ndarray'>
[2 2 3 3 4 4]

Now you can use to_categorical():

from keras.utils import to_categorical
cat1 = to_categorical(npa)
print(cat1.dtype, type(cat1))
print(cat1)

Which yields a matrix:

float32 <class 'numpy.ndarray'>
[[0. 0. 1. 0. 0.]
 [0. 0. 1. 0. 0.]
 [0. 0. 0. 1. 0.]
 [0. 0. 0. 1. 0.]
 [0. 0. 0. 0. 1.]
 [0. 0. 0. 0. 1.]]

Note that the matrix contains five columns (starting at zero up to four, which is my max. value in the np.array). The first two columns (representing 0 and 1 in the original data) are 0 in the whole matrix, because none of these values are found in the original data.

to_categorical() also takes input which is not explicitly defined as np.array. For instance the statements below would also be legal.

alt1 = to_categorical([0,0,1,1,2,2])
print(alt1.dtype, type(alt1))
print(alt1)

alt2 = to_categorical((0,0,1,1,2,2))
print(alt2.dtype, type(alt2))
print(alt2)

Because the range of values now is between 0 and 2, the result would look like:

[[1. 0. 0.]
 [1. 0. 0.]
 [0. 1. 0.]
 [0. 1. 0.]
 [0. 0. 1.]
 [0. 0. 1.]]

Pandas: get_dummies()

When you have a Pandas df, you can convert some column to dummies using get_dummies(), regardless of the data type in the column. So it is also possible to convert a column of strings to dummies.

import pandas as pd
df = pd.DataFrame(data={'col1':["A", "A", "B", "B", "C", "C"]})
alt3 = pd.get_dummies(df['col1'])
print(type(alt3))

This gives:

<class 'pandas.core.frame.DataFrame'>
   A  B  C
0  1  0  0
1  1  0  0
2  0  1  0
3  0  1  0
4  0  0  1
5  0  0  1

Note that the result is (again) a Pandas df. So we need to convert it to a np.array.

alt3 = alt3.to_numpy()
print(alt3.dtype, type(alt3))
print(alt3)

This yields:

uint8 <class 'numpy.ndarray'>
[[1 0 0]
 [1 0 0]
 [0 1 0]
 [0 1 0]
 [0 0 1]
 [0 0 1]]

So that it is ready to be used with Keras.

Note that the matrix generated here does not (!) start at zero. Instead each distinct value in the chosen Pandas column gets it's own column in the dummy matrix.

2 of 2
0

Try:

X = dataset[:,0:17].astype(float).astype(int)

I think that if you have a string like '45.2', you will have to cast it as a floating-point first and from float you can cast them into integer.

I will be glad if an editor could corroborate/correct this answer.

🌐
Educative
educative.io › answers › how-to-perform-one-hot-encoding-using-keras
How to perform one-hot encoding using Keras
The Keras API provides a to_categorical() method that can be used to one-hot encode integer data.
🌐
Machinecurve
machinecurve.com › index.php › 2020 › 11 › 24 › one-hot-encoding-for-machine-learning-with-tensorflow-and-keras
One-Hot Encoding for Machine Learning with TensorFlow 2.0 and Keras | MachineCurve.com
If we need to convert our dataset into categorical format (and hence one-hot encoded format), we can do so using Scikit-learn's OneHotEncoder module. However, TensorFlow also offers its own implementation: tensorflow.keras.utils.to_categorical.
🌐
Victorzhou
victorzhou.com › blog › one-hot
One-Hot Encoding, Explained - victorzhou.com
Below are several different ways to implement one-hot encoding in Python. Using scikit-learn’s OneHotEncoder: from sklearn.preprocessing import OneHotEncoder encoder = OneHotEncoder(sparse=False) print(encoder.fit_transform([['red'], ['green'], ['blue']])) ''' [[0. 0. 1.] [0. 1. 0.] [1. 0. 0.]] ''' Using Keras’s to_categorical: from keras.utils import to_categorical print(to_categorical([0, 1, 2])) ''' [[1.
Find elsewhere
🌐
Kaggle
kaggle.com › code › selcukcan › nlp-6c-onehot-encoding-and-word-embedding-in-keras
NLP 6c OneHot Encoding and Word Embedding in Keras | Kaggle
October 24, 2025 - Explore and run AI code with Kaggle Notebooks | Using data from No attached data sources
🌐
GitHub
github.com › christianversloot › machine-learning-articles › blob › main › one-hot-encoding-for-machine-learning-with-tensorflow-and-keras.md
machine-learning-articles/one-hot-encoding-for-machine-learning-with-tensorflow-and-keras.md at main · christianversloot/machine-learning-articles
November 24, 2020 - If we need to convert our dataset into categorical format (and hence one-hot encoded format), we can do so using Scikit-learn's OneHotEncoder module. However, TensorFlow also offers its own implementation: tensorflow.keras.utils.to_categorical.
Author: christianversloot
🌐
Quora
quora.com › How-can-I-use-one-hot-encoding-as-input-for-the-input-layer-in-Keras
How to use one-hot encoding as input for the input layer in Keras - Quora
Answer (1 of 3): If your inputs are already in one-hot format you don’t need anything else. Just plug-and-play. If you need to convert them first one_hot or to_categorical may help.
🌐
Google
developers.google.com › machine learning › machine learning glossary
Machine Learning Glossary | Google for Developers
The encoder maps the input to a (typically) lossy lower-dimensional (intermediate) format.
🌐
Google
developers.google.com › machine learning › prerequisites and prework
Prerequisites and prework | Machine Learning | Google for Developers
While focusing on core ML concepts, the course incorporates practical programming exercises using libraries like NumPy, pandas, and Keras but doesn't delve deep into specific ML APIs.
🌐
Medium
medium.com › @sanjay_dutta › what-is-one-hot-encoding-in-machine-learning-a-comprehensive-guide-with-examples-090f037a6bbe
What Is One-Hot Encoding in Machine Learning? A Comprehensive Guide with Examples | by Sanjay Dutta, PhD | Medium
December 1, 2024 - One-hot encoding is a method for converting categorical data into a binary vector representation. Each category is represented by a vector with a length equal to the total number of unique categories.
🌐
Coursera
coursera.org › home › categories › data science › machine learning
IBM AI Engineering Professional Certificate | Coursera
Explain how one-hot encoding, bag-of-words, embeddings, and embedding bags transform text into numerical features for NLP models
Rating: 4.6 ​ - ​ 22.3K votes
🌐
Kaggle
kaggle.com › code › praneet460 › one-hot-encoding
One Hot Encoding
February 15, 2019 - Explore and run AI code with Kaggle Notebooks | Using data from No attached data sources
🌐
GeeksforGeeks
geeksforgeeks.org › nlp › one-hot-encoding-in-nlp
One-Hot Encoding in NLP - GeeksforGeeks
In one-hot encoding each unique word is mapped to a binary vector where only one position has the value 1 and all others are 0.
Published: January 7, 2026
🌐
Facebook
facebook.com › groups › 1892701891124318 › posts › 2946822575712239
Eps: 1 ONE HEART (Satu Hati) maaf jarang Up - -
Popular groups · Find communities for you · Over 1 billion people across the globe are using Facebook Groups to explore their favorite topics · Log in · Categories · Science & tech · Travel · Animals · Sports & fitness · Entertainment