Why Are There So Many Tamil Encodings?
Before Unicode became the global standard in the 2000s, every software vendor created their own way of encoding Tamil characters. Since computers originally stored only 128 or 256 characters (the ASCII set), Tamil font vendors mapped Tamil letters to those existing slots. Each vendor chose different mappings, resulting in dozens of incompatible encodings. A document created in one encoding is unreadable without the exact same font installed.
Understanding these encodings helps you correctly identify and convert legacy Tamil documents.
The Major Tamil Legacy Encodings
Bamini
Bamini is the most widely used Tamil font encoding in India. Developed by Murasu Systems, it was adopted by Tamil Nadu's print media in the late 1980s and became the dominant standard for Tamil newspapers, magazines, and books throughout the 1990s and 2000s. Bamini text stored as ASCII looks like a random mix of English letters — for example, the Tamil word "வணக்கம்" (vanakkam) appears as "tzkf;fk;" in raw Bamini encoding.
Common uses: Tamil Nadu newspapers, Eelam Tamil publications, government documents, early Tamil books.
TSCII (Tamil Script Code for Information Interchange)
TSCII was an attempt in the late 1990s to create a standardised Tamil encoding before Unicode became mainstream. It was designed by a committee of Tamil computing organisations and was widely adopted for Tamil internet content in the early 2000s. TSCII uses the upper 128 ASCII slots (values 128–255) to store Tamil characters, making TSCII files slightly more structured than Bamini.
Common uses: Early Tamil websites, Tamil Nadu government digital archives, TSCII-encoded email lists.
TAM / TamShakti
TAM (also called TamShakti or Tam Shakti) is another legacy encoding that was popular in the diaspora Tamil community, particularly in Sri Lanka, Malaysia, Singapore, and the UK. It has a different character mapping from both Bamini and TSCII.
Common uses: Sri Lankan Tamil media, Malaysian Tamil publications, diaspora Tamil community sites.
SenTamil
SenTamil is a font encoding developed for early Tamil desktop publishing and was widely used alongside Bamini. Many publishers switched between Bamini and SenTamil depending on the software they used, so you will often find SenTamil-encoded files in Tamil Nadu publishing archives.
Common uses: Tamil Nadu publishing houses, literary magazines.
Vanavil
Vanavil fonts were developed for use with the Vanavil Tamil word processor, one of the most popular Tamil DTP applications of the 1990s. Vanavil has a distinct encoding and is often found in content produced by Tamil Nadu state government departments that used the Vanavil software.
Common uses: Tamil Nadu state government documents, Vanavil word processor files.
SunTommy and SunSuraj
The Sun fonts (SunTommy, SunSuraj, SunMadurai, etc.) were developed by SunFonts for use in Tamil publishing. They were popular in South India and are often found in Tamil text books and educational materials from the 2000s.
Shree Lipi
Shree Lipi is a multi-script DTP software and font system developed by Modular InfoTech. It supports Tamil along with many other Indian scripts. Shree Lipi Tamil encoding is commonly found in formally printed Tamil content from Maharashtra-based publishers.
MCL Mangai
MCL fonts were developed by Micro Computer Lab and are commonly used in Tamil educational publishing, particularly for textbooks and academic materials.
How to Identify Which Encoding a Document Uses
There are three reliable ways to identify the encoding:
- Check the font name in the source document. Open the file in Microsoft Word, select the Tamil text, and look at the font name in the toolbar. The font name usually matches the encoding name directly (e.g., font "Bamini" = Bamini encoding).
- Try each encoding in the converter. Paste the text into FontConvertor.in, try Bamini first, then SenTamil, then TSCII. The output will look correct when you have the right encoding selected.
- Look at the raw characters. Bamini text tends to use semicolons, colons, and common English punctuation heavily. TSCII text uses characters in the 128–255 range which appear as accented Latin characters (é, ê, etc.) on some systems.
Unicode: The Final Destination
All of these legacy encodings should ideally be converted to Unicode (UTF-8) for long-term preservation. Unicode Tamil is standardised, universally supported, and searchable. FontConvertor.in supports all 18+ Tamil encodings listed above and converts between any of them in real time.