scieee AI-readable full text Open interactive document viewer

Suganita

Chawla, Aman

Abstract

Part I focuses on the lexical layer---the fundamental vocabulary of the language. It establishes design goals, outlines a phased development plan, and provides a complete token specification for versions v0--v2. The token tables cover both high-level language keywords (function definitions, control flow) and low-level virtual machine operations (stack manipulation, arithmetic, hardware I/O). Each token is documented with its Devanāgarī form, ASCII transliteration, and English gloss linking it to its philosophical background

Full text

Suganita Part I: Token Specification A. Chawla REAL Institute December 3, 2025 Abstract This document presents the first part of a multistage blueprint for �(``well-computed''), a programming language written entirely in Devanāgarī script and designed for microcontroller platforms, specifically the Arduino Uno (ATmega328P). Unlike superficiallocalization efforts that simply translateWestern programming constructs, � draws conceptually from Indian intellectual traditions: Vedic mathematics, Nyāya logic (classical Indian reasoning), and Pāninian grammatical theory. Part I focuses on the lexical layer---the fundamental vocabulary of the language. It establishes design goals, outlines a phased development plan, and provides a complete token specification for versions v0--v2. The token tables cover both high-level language keywords (function definitions, control flow) and low-level virtual machine operations (stack manipulation, arithmetic, hardware I/O). Each token is documented with its Devana gariform, ASCII transliteration, and English gloss linking it to its philosophical background. Key design choices include representing ``no operation'' as  (śūnya, constructive pause rather than meaningless gap) and modeling conditional jumps after the structure of classical Indian argumentation. The hardware constraints of the Arduino Uno (2 KB RAM, 32 KB flash, 8-bit architecture) enforce discipline on the tokendesign, ensuring the languageremainsimplementableonresource-constrained platforms while preserving its conceptual foundations. This work is dedicated to Kriya Yoga master Paramahamsa Hariharananda Giri (1907-2002). 1 Contents 1 Design goals for � 3 1.1 Motivationandidentity.............................. 3 1.2 Hardware and engineering constraints . . . . . . . . . . . . . . . . . . . . . . 3 1.3 Philosophical and mathematical background . . . . . . . . . . . . . . . . . . . 4 1.4 ScopeofPartI .................................. 4 2 Development plan overview 6 2.1 Phase 0: Concept and constraints . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.2 Phase 1: Arduino–hosted VM prototype . . . . . . . . . . . . . . . . . . . . . 6 2.3 Phase 2: External compiler and language syntax . . . . . . . . . . . . . . . . . 7 2.4 Phase 3: Assembly implementation of the VM . . . . . . . . . . . . . . . . . . 7 2.5 Phase 4: Lightweight ML augmentation . . . . . . . . . . . . . . . . . . . . . 7 2.6 Phase 5: Self–hosting and expansion . . . . . . . . . . . . . . . . . . . . . . . 8 3 Token specification for versions v0–v2 9 3.1 Corelanguagekeywords ............................. 9 3.2 Control flow and logical structure . . . . . . . . . . . . . . . . . . . . . . . . 10 3.3 Operators and punctuation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 3.4 Virtual machine opcodes (v0–v2) . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.5 Built–in hardware functions (source–level) . . . . . . . . . . . . . . . . . . . . 14 3.6 ML–related tokens (planned for v2) . . . . . . . . . . . . . . . . . . . . . . . 15 4 Summary and outlook 17 2 1 Design goals for � 1.1 Motivation and identity The language � is intended to differ from a purely transliterated or ``localised'' Western language in three crucial ways: 1. Script and surface form. All identifiers, keywords, and structural tokens of the language are expressed in Devana gari. Source files are stored as UTF--8 and are visually legible to anyone familiar with Sanskrit or modern Indian languages using this script. 2. Conceptual grounding. The interpretation of tokens, especially at the control, arithmeticand logicallevels, isguidedbyIndianmathematical andphilosophical categories. For example, the instruction traditionally called ``NOP'' (no operation) is replaced by  (śūnya), representing a constructive pause rather than a meaningless gap. 3. Hardware orientation. The first several years of the project explicitly target the Arduino Uno. This constraint enforces clarity, compactness, and discipline on the language design: the token set must be implementable in a few kilobytes of RAM and a few tens of kilobytes of flash, on an 8–bit architecture. 1.2 Hardware and engineering constraints The initial target architecture is the ATmega328P microcontroller, as configured on the Arduino Uno board: •8–bit AVR core at 16 MHz •2 KB SRAM •32 KB flash program memory (approximately 30 KB usable after bootloader) •1 KB EEPROM •No hardware floating point unit •Simple peripherals: timers, ADC, PWM, digital IO, UART 3 These constraints inform the token design: •The virtual machine will be stack–based, using 16–bit logical words implemented over 8–bit registers. •The bytecode instruction space is limited; opcodes must be compact and semantically dense. •Strings and Devana garitext will be emitted as UTF–8 sequences but interpreted as opaque bytes by the VM. •Arithmetic in v0–v2 will focus on integers; fixed–point or scaled integer arithmetic may later represent real values. 1.3 Philosophical and mathematical background At a conceptual level, � aligns core language affordances with six strands of the Indian knowledge tradition: Vedic mathematics informs arithmetic semantics; Nyāya logic shapes control flow; Pāninian grammar guides lexical discipline; Sānkhyaand Yoga concepts govern executionstate; Vākya tradition ensures meaningful surface forms; and embedded ML praxis enables lightweight inference. The goal is not mere ornamentation, but principled mapping from philosophical categories to concrete implementation. Concept diagram (overview). The diagram below sketches how traditions inform layers of the language. It is illustrative rather than prescriptive. Pragmatic constraints. Allmappings must respect the ATmega328Penvelope: 2KB SRAM, ∼30KB usable flash, 8-bit arithmetic, and no FPU. Accordingly, v0--v2 prioritize: integer-first arithmetic, compact branching encodings (16-bit jump targets), UTF-8 treated opaquely in the VM, and ML as minimal fixed-point primitives callable without complicating the core interpreter loop 1.4 Scope of Part I This Part I addresses the lexical layer only: •Keyword and operator names 4 Vedic mathematics Nyāya logic Pāninian grammar Sāṅkhya & Yoga Vākya surface Embedded ML praxis Surface tokens Controlflow semantics Arithmetic semantics VM state & timing ML hooks & fixedpoint �, , , , � �, ,  (JZ),  (JNZ)  (MUL),  (DIV),  (MOD) � (HALT),  (NOP), � (DELAY) , �, _, _,  Figure 1: Conceptual mapping from Indian traditions to language layers and token examples in � •Punctuation tokens and structural delimiters •Virtual machine opcode names and their semantic labels •Built–in function names for hardware interaction Grammar, semantics, and virtual machine implementation details will appear in later parts. 5 2 Development plan overview The overall development path is structured into phases. Only the tokens relevant to v0--v2 are specified in this document; however, the plan contextualizes how these will be used. 2.1 Phase 0: Concept and constraints •Fix the hardware target to Arduino Uno / ATmega328P. •Decide on UTF--8 as the sole source encoding for .su files. •Choose a stack–based virtual machine model and a 16–bit logical word size. •Reserve ranges of bytecode opcodes for arithmetic, control, IO, and future ML extensions. •Decide that early implementation will be in Arduino C++ (for ease), followed by a rewrite in AVR assembly. 2.2 Phase 1: Arduino–hosted VM prototype •Implement a minimal interpreter loop in C++ on the Arduino Uno. •Support: –stack operations, –integer arithmetic, –unconditional and conditional jumps, –serial output, –a HALT instruction. •Hard--code small bytecode pages in flash to test the VM. •Use simple Devana garistrings for serial output to verify the complete UTF--8 path. 6 2.3 Phase 2: External compiler and language syntax •Implementa compiler on a hostcomputer (language still to bechosen) which: –reads � source code, –tokenizes using the token specification in this document, –parses according to the basic grammar (to be specified later), –emits bytecode (.subc) using the opcode set defined here. •Introducelanguage--level keywords (functiondefinitions, conditionals, loops) mapped to the low–level VM tokens. 2.4 Phase 3: Assembly implementation of the VM •Rewrite the VM core in AVR assembly, using the same opcode semantics. •Allocate fixed memory regions for: –VM state (instruction pointer, stack pointer, frame pointer), –data stack, –global variables, and –hardware state caches if needed. •Implement the hardware IO opcodes: digital IO, analog input, delay, and serial. 2.5 Phase 4: Lightweight ML augmentation •Introduce ML tokens for simple models suitable for an 8–bit microcontroller: –linear models or very small multi–layer perceptrons with a handful of parameters, –hand–coded fixed–point inference. •Ensure the ML tokens are integrated without altering the core VM structure. 7 2.6 Phase 5: Self–hosting and expansion As the toolchain stabilizes: •Gradually rewrite the host compiler in � itself. •Port the VM and language to more powerful microcontrollers. •Expand the token set conservatively when new capabilities are needed. 8 3 Token specification for versions v0–v2 The token system for � can be grouped into several categories: 1. Core language keywords 2. Types and literals 3. Operators and punctuation 4. Control–flow and structure keywords 5. Virtual machine opcodes 6. Built–in hardware functions 7. ML–related high–level constructs In the tables below, each row lists: •the Western or conventional concept, •the intended Devana garitoken, •an ASCII transliteration, •an English gloss. Only tokens planned for implementation in v0, v1 or v2 are included. 3.1 Core language keywords These are high–level language keywords, visible to the programmer in � source files. Western concept Devana garitoken Transliteration English gloss / role Function definition � ka rya Introduces a function definition 9 Threshold  sima Apply threshold (e.g. step function) These tokens are sufficient to encode simple control policies and decision rules learned elsewhere and embedded on the Arduino as constant parameters. 16 4 Summary and outlook This Part I specification has assembled: •a concise statement of overall design goals for �, •a phased development plan focused initially on Arduino Uno, •and a concrete, implementable token set for language keywords, operators, opcodes, hardware primitives, and basic ML hooks for versions v0--v2. Three design choices illustrate this conceptual alignment. First,  replaces a neutral "NOP" with a meaningful sunya operation. Second, control-flow operations like conditional jump are associated with  (cause) and  (example), echoing Nyāya structure. Third, � (cessation) gives semantic depth to the idea of halting computation. In Part II, ``Grammar and structural forms'', the next steps will be: 1. to define the precise lexical rules for identifiers, including permissible Devana garicodepoints and combining marks; 2. todescribe the context--freegrammarforexpressions, statements, andmodule structure; 3. to specify how control flow constructs map systematically to the underlying opcodes; 4. and to demonstrate small complete programs in � and their compiled bytecode. Subsequent parts will address the virtual machine layout, the Arduino C++ implementation, and the eventual AVR assembly version. Over time, as the language stabilizes, the token set may grow; however, the aim will be to preserve the conceptual clarity and Vedic orientation recorded in this first blueprint. Acknowledgments This work was produced with the assistance of large language models. 17 Tradition Core idea Language fea ture layer Example tokens Implementation notes Vedic mathemat ics Sūtradriven computa tion (e.g., ����: vertically and crosswise) Arithmetic se mantics and future specialized ops  (OP_MUL),  (OP_DIV),  (OP_MOD) Start with integer ops (v0–v2), reserve opcode space for patterned arithmetic; validate cost on 8bit AVR before inclusion. Nyāya logic Fivepart infer ence (�, , , , �) Control flow semantics and branching vocab ulary � (if),  (else);  (OP_JZ),  (OP_JNZ) Model conditional jumps as cause/ex ample; later sugar for multiclaim branching; keep jump targets 16 bit for compact code. Pāṇinian grammar Rulegoverned transformations (, �) Lexical discipline and parsing trans forms  (I as statement end), �� (: association),  (OP_CALL), � (OP_RET) Deterministic tokenizer over Devanāgarī; UTF8 treated as opaque bytes in VM; grammarlevel trans forms constrained to lineartime passes. Sāṅkhya and Yoga Purusha/Prakṛti distinction; � (cessation) Execution state and halting se mantics � (OP_HALT),  (OP_SHU;  pause) HALT maps to a quies cent VM state; NOP as constructive pause for timing/alignment; ex pose timing semantics via � (OP_DELAY_MS). Vākya tradition (śāstric clarity) Meaningful sur face aligned with roles Surface tokens and keyword design � (function),  (entry),  (type),  (string), � (void) Prefer semantically grounded lexemes over transliteration; maintain compact en codings; provide ASCII aliases only for early tooling. Early ML praxis (em bedded) Small, fixed mod els; affine maps and thresholds Highlevel ML hooks and com pact tensors , �, _, _, _��,  Hostside training; deviceside fixedpoint inference; weights in flash; preserve VM simplicity by treating ML as callable primi tives. Table 1: Mapping philosophical traditions to language and VM features (v0–v2 focus, with reserved room for later expansion). 18