Update:

Used indian names instead of prev dataset

result:

makemore

makemore takes one text file as input, where each line is assumed to be one training thing, and generates more things like it. Under the hood, it is an autoregressive character-level language model, with a wide choice of models from bigrams all the way to a Transformer (exactly as seen in GPT). For example, we can feed it a database of names, and makemore will generate cool baby name ideas that all sound name-like, but are not already existing names. Or if we feed it a database of company names then we can generate new ideas for a name of a company. Or we can just feed it valid scrabble words and generate english-like babble.

This is not meant to be too heavyweight library with a billion switches and knobs. It is one hackable file, and is mostly intended for educational purposes. PyTorch is the only requirement.

Current implementation follows a few key papers:

Bigram (one character predicts the next one with a lookup table of counts)
MLP, following Bengio et al. 2003
CNN, following DeepMind WaveNet 2016 (in progress...)
RNN, following Mikolov et al. 2010
LSTM, following Graves et al. 2014
GRU, following Kyunghyun Cho et al. 2014
Transformer, following Vaswani et al. 2017

Usage

The included names.txt dataset, as an example, has the most common 32K names takes from ssa.gov for the year 2018. It looks like:

emma
olivia
ava
isabella
sophia
charlotte
...

Let's point the script at it:

$ python makemore.py -i names.txt -o names --device mps

Training progress and logs and model will all be saved to the working directory names. The default model is a super tiny 200K param transformer; Many more training configurations are available - see the argparse and read the code. Training does not require any special hardware, it runs on my Macbook Air and will run on anything else, but if you have a GPU then training will fly faster. As training progresses the script will print some samples throughout. However, if you'd like to sample manually, you can use the --sample-only flag, e.g. in a separate terminal do:

$ python makemore.py -i names.txt -o names --sample-only

This will load the best model so far and print more samples on demand. Here are some unique baby names that get eventually generated from current default settings (test logprob of ~1.92, though much lower logprobs are achievable with some hyperparameter tuning):

Vigneshwaran
Kriti
Arpita
Indora
Aqsa
Anwar
Apratim
Hasni
Dalia
Aadharsh
Amarnath
Bramoni
Nethaji
Chetna
Aastha
Aakriti
Abhyuday
Akarsh
Navneet
Om
Swapnil
Akhil
Madhav
Divyam
Kirti
Annanya
Arunima
Niteesh
Nithisha
Sandhya
Uttam
Vatsalya
Aparna
Disha
Seema
Suryansh
Yashaswi
Avinash
Smita
Subhra
Tuhin
Arti
Aakash
Ayan
Aryan
Vertika
Umme
Yukti
Vishesh
Ayushman
Geeta
Divyanshu
Jaswant
Hemanshu
Bhoopendra
Bhawana
Divakar
Dimpal
Shrestha
Shivanshi
Shubh
Shreyash
Satyarth
Satyam
Shilpa
Shikha
Soniya
Tanay
Sushant
Ujjwal
Tanu
Subham
Srishti
Suryakant
Sudhanshu
Shivang
Shilpi
Sourabh
Sonu
Shakti
Satendra
Shaqib
Shalu
Sudheer
Vipul
Vipin
Yashika
Virat
Varsha
Sunny
Vineeta
Vibhanshu
Monti
Minakshi
Nitikarsh
Mugdha
Anukriti
Kritika
Jaya

Have fun!

License

MIT

Name		Name	Last commit message	Last commit date
Latest commit History 25 Commits
names		names
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
makemore.py		makemore.py
names.txt		names.txt
res.png		res.png
res2.png		res2.png

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Update:

result:

makemore

Usage

License

About

Releases

Packages

Languages

License

jayeshthk/makemore

Folders and files

Latest commit

History

Repository files navigation

Update:

result:

makemore

Usage

License

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages