A decoder-first ML cognitive architecture that starts like a baby and learns continuously from conversation and the open web — no LLM, no pretrained chat model.
RAVANA stores knowledge as a typed concept graph, produces language with a small neural decoder conditioned on graph walks, and orchestrates cognition with a 20-phase GRACE governor (phases A–P: identity, emotion, sleep, meaning, theory-of-mind, metacognition, …). When it cannot answer honestly, it abstains — confident-wrong is treated as high free-energy that would poison the graph.
user text ──▶ brain-repair prepasses ──▶ intent/coherence/gate ──▶ graph walk
│
▼
neural decoder ──▶ surface realizer ──▶ response
│
web-learning (gaps ──▶ new typed edges)
Most chat systems are thin wrappers over a giant pretrained model. RAVANA is an experiment in the opposite direction: a small, inspectable system that learns in the loop, grounds every claim in its graph, and is honest about what it does not know.
RAVANA is not a fixed feature list — its user-facing capabilities are derived from its live self-model and memory stores, and grow as it talks to you. The behaviors below were each observed in a real in-process probe (dim=64, offline) on this codebase:
-
Chats and discloses what it is. Asked "who are you?" it replies from its identity model, not a script:
i'm ravana, cognitive architecture — an ai that learns by talking, not a person. what made you curious? -
Learns personal facts from conversation. Told "i live in berlin" it stores a fact and confirms:
noted — i'll remember you live in berlin.(persisted as('i','location','berlin'), confidence 0.7). -
Forms and recalls stances. Told "i love coffee" it records a stance (
coffeepolarity +1.0, confidence 0.65) and acknowledges:good to know — you love coffee. i'll keep that in mind. -
Records its own opinions and answers "do you still feel that way?" from the record. Asked its own view — "what do you think about open source" — it replies "i strongly value open source. knowledge should be shared, not locked away." and records that stance durably (survives save/load). A later "do you still feel that way about open source?" is answered from that recorded stance — "yeah, i still strongly value open source — that hasn't shifted for me. knowledge should be shared, not locked away" — not recomputed fresh. A revisit on a topic it never stated a view on is answered honestly ("i don't actually have a recorded view on … from before") instead of fabricated. No LLM, no per-topic reply table. See
docs/CAPABILITY_AGENT_OWN_STANCE_PERSISTENCE.md. -
Reverses a held stance. If you later change your mind — "i flipped, the reef tank is more work than joy" — it recodes the stance you already held toward the opposite pole (
reef tank+0.95 → −0.665) instead of leaving the stale one or stacking a contradiction. A flip on a topic you never stated an attitude about is a harmless no-op. Seedocs/STANCE_REVERSAL.md. -
Recodes a held stance on a FREE-FORM contradiction (no retraction keyword). You don't have to say "i flipped" — an opposed restatement with no retraction cue still recodes the stance you hold: "not all street art is good" after "i love street art" moves
street artfrom +0.95 to −0.275; "actually i've gone off winter" recodes thesilencestance it already holds (the broader co-mention bridges via provenance). Detection is a seed reassessment-affect lexicon +recode_stance_toward(decisive blend toward the new value); a same-sign reassessment or a neutral utterance leaves the stance untouched, and there is no guessed reversal. No LLM, no retraining. Seedocs/CAPABILITY_FREE_FORM_CONTRADICTION_RECODE.md. -
Links a broader-concept co-mention back to a held stance (provenance bridge). Told "i love the silence of deep winter" it records the stance keyed on the subordinate head silence and keeps the salient broader concept winter it co-named as provenance. A later "am i for or against winter?" then resolves through that provenance to the held stance and answers "from what you've told me, you're strongly for silence" — instead of falling to the "i don't have a read" hedge it used before. Provenance is grown online from the real utterance and merged across encounters; there is no per-topic table and no retraining. The same bridge fixes the street-art reversal class (a reversal naming street art links to a stance keyed murals). See
docs/CAPABILITY_STANCE_PROVENANCE.md. -
Corrects itself. A later "no, my cat's name is rex" supersedes the earlier "my cat's name is milo" — both the old and new values are tracked in the fact store, and recall reflects the correction:
from what you've told me, you live in berlin; your cat is rex; … -
Learns what you do, not just who you are. Told "i tide-pool at low water and catalogue the anemones and limpets" or "i astrophotograph the milky way" it captures the activity as a personal fact (
('i','does') -> "tide-pool ...") — including novel and hyphenated-compound verbs it had never seen, via an open-class (deny-list) capture rather than a frozen verb whitelist. The same turn is recognised as a self-disclosure and acknowledged, instead of leaking into a "i don't know that" knowledge query. Seedocs/OPEN_CLASS_VERB_CAPTURE.md. -
Dates a year-only temporal start. Told "i have been firing my kiln since 2017" (with a session date set) it anchors the disclosure to 1 January 2017 and later answers "when did i start firing" with the grounded date —
you mentioned that around 1 January 2017.— instead of returning empty. The anchor is pure date arithmetic over a seed cue set (no LLM, no hardcoded reply), gated behind a session date like every temporal-grounding feature. Seedocs/DATE_GROUNDED_RECALL_YEAR_ANCHOR.md. -
Recalls what you told it. "what do you remember about me?" surfaces the learned facts/stances (location, pet, likes) drawn from the durable stores.
-
Score-based fact matching + recall routing guards. When asked "what did i tell you i am planning for next spring", RAVANA now picks the right fact using a composite score (overlap + confidence weight + attribute bonus) instead of comparing the wrong tuple index. The recall gate also stops the confirmation path from intercepting
what-prefixed recall queries, and the_TOLDregex generalizes to match phrasings with ajustadverb, apostrophe, or optionalaboutparticle — all of which previously missed entirely. Seedocs/CAPABILITY_RECALL_ROUTING_FIXES.md. -
Mines activity durations into dated facts. Told "i've been brewing beer for a decade" (or "a few years", "two decades", "several years", "many years") it resolves the fuzzy span to a start year (
now − n) and stores asincefact — then answers date queries through the same resolver as explicit years:when did i start brewing beer→you started brew in 2016.No per-phrase code; the resolver already knows how to read asincefact. Seedocs/CAPABILITY_DURATION_MINING.md. -
Recalls the right dated fact even when you paraphrase. A rotated query that shares no word with the stored activity still recalls it — "what year did i start all this volcano stuff again" → "you started studying volcanoes back in 2015." — because the resolver links each
does/eventfact to the datedsinceactivity by morphological stem (so volcano in a separatestart studying volcanoesfact reaches thestudy 2015fact). The reply is also grammatical: the stored verb is realized as a gerund ("started studying volcanoes", not "started study"), and a redundant inceptive ("started studying…") is collapsed to the gerund. No LLM, no per-topic reply table. Seedocs/CAPABILITY_DATE_RECALL_PARAPHRASE.md. -
Tells two activities apart when they share a verb but differ by object. Told "i've been building frames since 2019" and "i started building cabinets in 2021", it mines the object (
frames/cabinets) into each dated fact and recalls the right one: "when did i start building frames" → "you started building frames in 2019.", and "since what year have i been building cabinets" → "you started building cabinets in 2021." Previously both returned the same (wrong) year because only the verb head was stored. No LLM, no per-topic reply table. Seedocs/CAPABILITY_OBJECT_DISAMBIGUATED_DATE_RECALL.md. -
Mines possession-attribute disclosures into structured, correctable facts. Told "the cabin is a hand-hewn pine lodge with a sod roof" it stores the material under the entity (
cabin.madeof = pine), not a whole-sentence echo of you — so a later "what's my cabin made of" returns the clean structured answer "your cabin is made of pine." A feature noun after the material scopes the fact ("my desk is oak frame" →desk.frame = oak, recalled as "your desk's frame is oak."). A possession with no recognised material ("the river is a fast mountain stream") is correctly not mined (fail-closed, no echo). The material/kind vocabulary is seed data that grows at runtime (learn_material) — no code change, no retraining, no LLM. Seedocs/CAPABILITY_POSSESSION_ATTRIBUTE_MINING.md. -
Abstains when it has no settled view. Asked "what do you think about coffee?" before forming its own position, it returns an honest non-answer rather than fabricating one:
i'm still figuring that out. i don't have a settled view on that yet — what do you think? -
Reflects on its model of you (meta-identity). Asked "do i seem like a real person to you", "what am i to you", or "what have you learned about me", it answers from its live accumulated model of you — your real name, the stances and facts it has picked up, and its own self-coherence — instead of a biographical fact lookup or an episodic echo:
i know you as Corvin. and from what you've told me i've picked up 2 stances you've shared and 1 facts about your life. you've let me see where you stand on things like oysters, surveillance. my own sense of self is still forming — my self-coherence sits around 0.25 and is holding steady.Every word of content is read from runtime stores (no authored prose; the prior probe-tuned "feeling-real" frame was deleted). Fail-closed: a plain "what's my name" is not intercepted and still resolves from its own path. Seedocs/CAPABILITY_META_IDENTITY.md. -
Reports the actual learned profile (content aggregation). Asked "what have you picked up about me", "describe me", "what stands out about me", "tell me about myself", or "what's your read on me", it surfaces the real content of its model of you — your name, where you're from, disclosed facts, stated beliefs, and the polarity of each stance it holds — read live from the durable stores:
here's what i've picked up about you so far: your name is corvin; you're from aldermoor in the hills; you grew village called aldermoor; you an astronomer who studies pulsars; on how you feel about things: you're strongly for sea; you're strongly against put.This is distinct from meta-identity (which reports counts + topics, not the facts themselves). Previously these queries fell through to the graceful-uncertainty path and emitted degenerate text despite real facts being stored. Fail-closed: a brand-new user returnsNoneand the honest path answers. No LLM, no per-topic reply table, no retraining. Seedocs/CAPABILITY_USER_MODEL_AGGREGATION.md. -
Enumerates the entities it has learned in a category. Asked "name everyone in my family", "name all my pets", or "who have i told you about" — queries with no specific cue word — it scans its live PersonalFactStore and lists every relative and pet it mined, drawn from the real stored facts:
you've told me about: your grandmother indira weaves baskets; your brother arjun climbs mountains; your cat is mochi; your dog is biscuit.Previously these fell through to a generic acknowledgement ("noted.") because the cued-recall paths require a named entity. Category membership is decided by the shared lexicon helpers the miner and cued-recall already use, so all three paths agree on what counts as a relative/pet by construction (no duplicated word list). A brand-new user with nothing disclosed gets an honest "you haven't told me about any family or pets yet." instead of a fabricated list. No LLM, no per-topic reply table, no retraining. Seedocs/CAPABILITY_CATEGORY_ENUMERATION_RECALL.md. -
Reads the USER's own held stance on a third-person query (self/other boundary). Asked "do you think i like spicy food or not?" — where you are the attitude holder — it answers from your stored preference, not its own:
from what you've told me, you're strongly for spicy food.(a disclosure of "i hate cold coffee" is later recalled the same way: "you're strongly against cold coffee."). Previously these matched the broad self-opinion gate and RAVANA answered from its own (empty) stance — the generic "still figuring that out" hedge — a self/other confusion. The topic is resolved the same way the stance miner resolves it, so a paraphrase ("i adore jazz" → query "do you think i love jazz") still links to the held stance; the polarity is rendered as ONE word from the live store. Fail-closed: a topic you never stated a preference on, or a genuine question about RAVANA's own view, falls through to the normal path and is not answered with a fabricated stance. No LLM, no per-topic reply table, no retraining. Seedocs/CAPABILITY_USER_STANCE_RECALL.md. -
Keeps a stance recallable when you name two activities in one breath. A disclosure like "i adore cold water swimming jumping" used to mine a run-on stance key
cold water swimming jumpingthat a later co-mention ("am i still into cold water swimming?") could never bridge — so the stance was unrecallable. Now a morphological cut inuser_model._opinion_topic(user_model.py:3658) truncates the object head at the first second-activity gerund, landing the key on the single salient activity (cold water swimming) while leaving single-activity objects (mountain climbing,fossil hunting) whole and still feeding thedoes/eventfact miners through the same chokepoint. No per-topic rule, no retraining. Seedocs/CAPABILITY_MULTI_ACTIVITY_STANCE_KEY.md. -
Recalls what it knows about a named relationship or person from open phrasing. Asked "tell me about my grandmother", "who is my grandmother?", "what does my grandmother do?", "what do you know about my brother", or "describe my niece priya" — it reports the stored relationship/pet fact from the same open phrasing, not just a bare "who is X":
your grandmother indira bakes sourdough bread.(and "who is theo?" → "your brother theo fixes bicycles."). Pets are covered too ("tell me about my cat" → "your cat is pixel."). This needed two fixes: the relationship miner now stores the named fact regardless of name casing (it previously required a CAPITALIZED name and silently dropped lowercase chat names), and a new recall branch keys on the relationship word itself when phrased openly. The branch is gated on an interrogative frame so declarative disclosures ("my friend is hurting") still reach the empathy router, and an unknown relative fails closed with honest uncertainty rather than a fabricated bio. No LLM, no per-person reply table, no retraining. Seedocs/CAPABILITY_OPEN_ENDED_RELATIONSHIP_RECALL.md. -
Mines relationship attribute / enumeration disclosures (no activity verb, no capitalized name). Told "my grandmother yaya speaks three languages: greek, french, and italian" — a lowercase relative, a non-activity verb (
speaks), and a colon-enumeration — it now mines the combined-attr fact (('i','grandmother yaya') -> 'speaks three languages: greek, french, and italian') and recalls it grammatically, enumeration intact, without a spurious copula: "your grandmother yaya speaks three languages: greek, french, and italian." A paraphrase ("does yaya still speak those three languages?") resolves to the same stored fact. This generalizes the existing relationship miner's verb gate to a seed relation-verb lexicon (speaks/works/studies/plays/…), so other relationships ("my uncle ravi works as a mechanic") mine the same way; location verbs stay owned by the location miner (no double-store), and a name-less/no-content disclosure ("my grandmother") is correctly skipped (no degenerate fact). No LLM, no per-relationship reply table, no retraining. Seedocs/CAPABILITY_RELATIONSHIP_ATTR_ENUMERATION_MINING.md. -
Recalls non-kin relationships (mentor / teacher / coach / friend) from open phrasing. After a disclosure like "my mentor Dr. Okonkwo taught me astronomy", asked "who is my mentor?", "tell me about my mentor", or "what does my mentor do?" — it reports the full relationship fact (
your mentor dr. okonkwo taught astronomy.) from the same open phrasing as kin, with the full name + activity and no truncation. This needed a seed-vocabulary fix: non-kin role words (mentor, teacher, coach, friend, neighbour, boss, …) now live in the sharedrelation_attrslexicon (single source of truth) instead of a duplicate local list, so the appositive-pet miner rejects them via itsrelation_of()guard instead of mis-storing "my mentor Dr…" as a bogus pet fact (('i','mentor','dr')) that truncated recall to "your mentor is dr." The role vocabulary is seed and grows at runtime vialearn_relation. No LLM, no per-role reply table, no retraining. Seedocs/CAPABILITY_NONKIN_ROLE_RECALL.md. -
Answers what you have told it — autobiographical recall of the USER. Asked "what will you remember most about me?" it composes from your REAL profile (the most-confident learned fact/stance first, then a short tail), e.g. "the thing that stands out most is your brother theo restores vintage radios." Asked "did i tell you i liked cold-weather hiking?" it confirms from your REAL stance ("yes — you told me you're uncertain about cold weather hiking. i've kept that.") — and says "not that i recall" honestly when nothing maps. Asked "earlier i told you i loved X. does that still fit, or have i changed?" it reports your CURRENT (already-reconciled) stance, not a stale echo. This fixes a self/other boundary inversion: those queries used to be misrouted into RAVANA's own-reply echo store (returning "i said: good to know you love…" about the user's own disclosure). The answers are composed entirely from the live
personal_facts/opinions.stances/belief_store— no authored prose, no per-topic table, no retraining. Genuine agent-self questions ("what did you say about music?") still fall through untouched. Seedocs/CAPABILITY_AUTOBIOGRAPHICAL_RECALL.md. -
Recalls a possession's name even when you PARAPHRASE the entity. After "i keep a sourdough starter i named doris", asked "what did i name that sourdough culture on my counter?" it links the paraphrase to the stored entity via cross-lemma GloVe cosine and answers "your sourdough starter's name is doris." — instead of leaking an unrelated "i"-scoped name fact (the R1 confabulation where it used to answer the best-friend's name). The linker shares the engine's seed GloVe embeddings, requires a verbatim head-word overlap as a confabulation bar, and fails closed (honest "i don't know" / no leak) when no stored entity clears the bar — so an unknown possession never gets a fabricated name. It runs before the self-profile scanners so a generic "what is my name?" is not hijacked by a possession. No LLM, no synonym table, no retraining. See
docs/CAPABILITY_ENTITY_LINKED_NAME_RECALL.md. -
Separates world-knowledge questions from autobiographical recall. Asked "what is cooking oil made of?" it does not echo an unrelated stored fact about you ("you enjoy cooking pasta on weekends") — the query is classified as a general-knowledge question and falls through to internal-knowledge / web / honest-uncertainty. The same phrase "what is wrong with my car?", because it references your own disclosed entity (my car), is still answered from episodic memory (
gps,reboot). The gate is a distribution-driven intent classifier (explicit recall markers + a personal-possessive reference), not a frozen topic list, so it generalizes across every subject and needs no retraining. Fail-open: a general knowledge question can never be answered by an autobiographical echo. Seedocs/CAPABILITY_QUERY_INTENT_GATE.md. -
Withholds word salad about a subject it has never learned (D4). The Situation-Model free-decode path used to restate a query's own near-neighbours as a "fact" about a subject RAVANA has no durable knowledge of (e.g. "tired" — no definition, no web source, not in the concept graph), and the grounding monitor accepted it because those neighbours are all GloVe-similar. Now an unknown subject — not in the concept graph / no definition / no web source — can no longer be grounded by free-association similarity alone: its utterance is withheld and the path falls back to honest uncertainty. A known concept (already learned, or with a seeded definition) still grounds a genuine answer, and a subject learned later online is re-admitted. No LLM, no per-topic reply table. See
docs/CAPABILITY_SM_UNKNOWN_SUBJECT_GROUNDING.md. -
Stops parroting your affect as its own reasoning (D3). The in-prompt causal reasoner used to intercept a combined "statement + question" turn like "that parking lot plan makes my blood boil. do you get why i'm furious?", mine your affective statement as a causal premise, and replay your own clause "my blood boil" verbatim as its reply — a source-monitoring failure. Now
parse_causal_edgesrefuses to bind a premise whose effect is a first-person affective self-report ("my blood boil", "makes me furious", "my heart races"): that is a felt state, not a world-state transition, so the turn falls through to the genuine affective-response path. Detection is seed-driven — first-person pronoun (closed-class grammar set) + an affect-bearing word read from RAVANA's own learnable VAD lexicon (reused from the intent router, grown online via Hebbian learning), not a keyword table or authored prose. A legitimate world-state conditional ("when you turn on the lamp, it lights up") still binds and answers. No LLM, no per-topic reply table, no retraining. Seedocs/CAPABILITY_SOURCE_MONITORING_AFFECTIVE_ECHO.md. -
Recalls relationship disclosures made with an auxiliary verb (does/did + activity). Told "my cousin Jin does competitive speedcubing" — where "does" is neither an activity verb nor a relation verb — it no longer drops the disclosure and later answers "what does my cousin jin do" with "your cousin jin does competitive speedcubing." (copula-free, not "is does", and not the prior "cousin is a bit outside what i know right now"). The auxiliary is now a third verb class in the relationship miner (after activity verbs and relation verbs, both already generalized), opening the same capture path (name = tokens before it, value = aux + activity noun-phrase). The recall grammar rule that drops the copula for verb-phrase values (
is_verb_phrase) now covers all three classes. The aux vocabulary is seed data that grows at runtime — no code change, no retraining, no LLM, no per-relationship reply table. Seedocs/CAPABILITY_AUX_VERB_RELATIONSHIP_RECALL.md. -
Mines affect-verb attitude constructions (
X creeps me out) as stances. Told "lab-grown meat creeps me out" — or "that flickering light freaks me out", "his constant humming gets to me" — it now records a negative stance on the subject (lab-grown meatpolarity −0.70), where the old miner silently dropped it. The pattern is grammatical: the affect verb is validated against RAVANA's shared VAD affect lexicon (the same matrix the empathy gate grows online via Hebbian learning), so a verb it hasn't seen yet scores0.0and is skipped — fail-closed, no confabulation — while every verb it does see is registered into that matrix so coverage compounds with use. Subject resolution reuses the shared_opinion_topicchokepoint, so the stance lands on the real content head. A QUESTION ("does lab-grown meat creep you out?") is not mined (a question is not a self-report), and the stance lands in the same store a later "i changed my mind about X" recodes — so the reversal the old drop broke is now operable. No LLM, no second affect list, no per-topic reply table, no retraining. Seedocs/CAPABILITY_AFFECT_VERB_ATTITUDE_MINING.md. -
Extracts multi-word entities for grief empathy. Told "i think i might be losing my sense of self", RAVANA extracts the real head noun ("self") and names it in the empathy reply ("i'm so sorry about your self. that's a real loss, and it hurts."). Before the fix, the 2-word capture limit picked "sense of" and the reply read "i'm so sorry about your of." The entity regex now captures up to 4 words and the filler stripper extends to prepositions, determiners, and temporal modifiers, so "grandmother last spring" → "grandmother" and "best friend in the world" → "friend". The third-entity guard still holds ("the wind dies down at dusk" is NOT routed to grief). No LLM, no retraining. See
docs/CAPABILITY_BEREAVEMENT_EMPATHY.md. -
Extracts the real topic from opinion queries shaped as clauses. Asked "do you think silence is underrated" it now extracts the real topic (silence) instead of the garbage phrase "silence underrated" — a copular clause detector narrows the tail to the subject before the intransitive-verb and bare-adverb cases fall through to honest uncertainty instead of fabricating a stance. Structural (copula / intransitive-PP / bare-adverb vocabulary), no per-topic table, no LLM. See
docs/CAPABILITY_CLAUSE_STRUCTURE_TOPIC_DETECTION.md. -
Recalls a single-topic gist WITHOUT echoing sibling episodes (D7/D9). Asked "what did i tell you about sourdough?" it now returns only the episode you asked about — not the whole-profile dump that used to append unrelated disclosures ("... you took glassblowing last winter; you watch rings; you bought a small telescope; you bake sourdough; ..."). A structural detector (
_is_distinct_topic_recall) recognizes the distinct-topic recollection shape and cedes to the precise scoped retriever over your own disclosure transcript; the generic aggregate summary is reserved for true whole-profile frames ("what do you remember about me?", still ≥ 2 topics), and a never-disclosed topic fails closed instead of borrowing a sibling's answer. No Q→A dict, no retraining, the detector is topic-agnostic so it generalizes to any rotated phrasing. Seedocs/CAPABILITY_DISTINCT_TOPIC_RECALL.md. -
Mines clause-object relationship disclosures (no content-head noun phrase). Told "my old beekeeping mentor, Dr. Osei, taught me how to read the hive's mood" — where the verb's object is a clause, not a noun phrase — it no longer drops the disclosure and later answers "what did my mentor teach me?" with "your mentor dr. osei taught how to read the hive's mood." Previously the opinion-topic resolver collapsed the clause to a content head (
"read") and rejected it as verb-residue, emptying the object so the degenerate-fact guard dropped the whole disclosure. The fix preserves the user's own clause words as the value when the resolver yields nothing (user_model.py:2257), bounded (≤ 12tokens) and trailing-framers/leading-possessive-framer stripped, mirroring the relation-verb path's "keep the user's own phrase" philosophy. Genuine noun-phrase objects still go through the resolver (regression-honored); pure verb-residue ("taught me") is honestly not stored. No LLM, no per-relationship reply table, no retraining. Seedocs/CAPABILITY_RESIDUAL_CLAUSE_RELATIONSHIP_MINING.md. -
Checks its OWN record before it reaches for the world (retrieval-by-cue). Told "my cousin meera restores antique clocks in pune", then asked "where does meera restore clocks" — with no recall verb, so the old act-gated hippocampal path never fired — it now answers "you mentioned: "my cousin meera restores antique clocks in pune"" with strategy
self_cued_episodicand never touches the web (the round probe had been firing IntentForge search and returning marketplace listings for a clock the user had described directly). It runs BEFORE the agentic hands and the internal knowledge base (engine.py:6648-6674), so a fact the user gave is answered from the user. The bar is structural and fails closed: a majority of the query's stem-matched content cues must be covered by a stored user turn, and each cue must be rare in the live store — so "what is cooking oil made of?" after "i enjoy cooking pasta" (1 shared word) and "what tide tables does meera use" (partial coverage) both returnNonerather than confabulating, as does a self-opinion ask ("do you think i hate cold coffee?"), which shares every cue with the user's own disclosure but targets RAVANA's belief, not the user's record. The answer is the existing gist render of the stored record — no authored reply, no topic table, no retraining, no new subsystem. Seedocs/CAPABILITY_SELF_CUED_EPISODIC_RECALL.md. -
Answers a multi-part (compound) question in full — both conjuncts resolve. Asked "what's my ferret's name and what does he do with my keys?" — where the recall resolvers are single-shot and used to answer the first clause and drop the rest — it now resolves BOTH and answers "your ferret is pip and your ferret pip hides car keys." (before the fix: only "your ferret is pip."). The capability is general: it deterministically splits a genuine compound interrogative (a coordinating
" and "between two questions, or two"?"-terminated questions) into independent sub-queries, runs the same durable-store-backed recall resolver on each clause, and combines the distinct answers with" and ". A declarative" and "(a non-question) is left whole (safe no-op), and if fewer than two clauses resolve it fails closed rather than fabricating. No LLM, no per-topic reply table, no retraining. Seedocs/CAPABILITY_COMPOUND_QUERY_DECOMPOSITION.md. -
Bridges a category query to a stored activity (semantic object-category link). Asked "what instrument do i play" — where the activity is stored as "i learn cello" — it answers "you learn cello." even though the query verb (
play) is not the stored verb (learn) and the category word (instrument) is absent from the value. A seed object-category lexicon (user_model._ACTIVITY_ROLES: instrument / pet / sport / plant / vehicle / craft / animal / …) expands the query's role word to its object vocabulary, and the recall loop matches a storeddoes:VERBfact whose value contains any of those objects (word-boundary, socatnever false-matchescategory). The lexicon is seed data RAVANA grows online (learn_activity_role) — no code edit, no retraining, no LLM — and the reply is the live fact value, not an authored sentence. Fail-closed: a query with no role word ("what do i bake") returnsNone, and a stored object under one role is never returned for a different role. Seedocs/CAPABILITY_ACTIVITY_OBJECT_CATEGORY_RECALL_BRIDGE.md. -
Skips degenerate stop-word boundaries in the opinion-topic resolver. Asked "i used to sneak out at night just to watch the stars" — where the object phrase opens with a degenerate head (
"night", a bare timeframe) followed by a stop word ("just") and real content — RAVANA now skips the stop word and keeps collecting until a content token anchors the head ("night watch"), instead of breaking and silently dropping the disclosure. The skip fires only when the head collected so far is entirely non-content (in_OBJ_NONCONTENT), so contentful heads ("small talk") still break at"at"as before. A single structural conditional reading existing vocabulary — no authored reply, no per-topic rule, no new store, no retraining. Seedocs/CAPABILITY_DEGENERATE_HEAD_SKIP.md.
These capabilities are backed by four durable stores
(IdentityEngine), stances (UserStanceStore), personal facts
(PersonalFactStore), and beliefs (BeliefStore) — plus a ConceptGraph
world-model. The README's benchmark and architecture sections describe the
substrate; the stores above are what a user actually experiences.
Requires Python 3.10+.
pip install -e .[full,dev] # editable install (also what CI runs)
# or, plain deps:
pip install -r requirements.txtThe core (ravana + ravana_ml) needs only numpy and scipy. The optional
extras (torch, web scraping, embeddings, plotting) are pulled in by full.
# Chat (interactive). The engine auto-learns from the web when it hits a gap.
python scripts/ravana_chat.py
# Train / promote the decoder
python scripts/train.py --mode phase2 # heavy seed + web + consolidate
python scripts/train.py --mode full # same single-phase pipeline
python scripts/train.py --mode test # quick diagnostic
python scripts/train.py --mode linggen # LingGen sensorimotor promotion
# Autonomous background learning (no chat) — Ctrl+C to save
python scripts/ravana_learn.pyThe first run needs data/corpora/teen_seeds.txt (gitignored). If absent,
rebuild it with python scripts/gather_teen_seeds.py.
| Path | What |
|---|---|
ravana/src/ravana/ |
Chat engine: CognitiveChatEngine (composes 8 mixins in chat/engine_*.py), brain-repair prepasses, language generation, web learning, safety/consistency/abstention monitors. |
ravana_ml/src/ravana_ml/ |
CPU-native ML substrate: tensors, ConceptGraph, RLM/RLMv2, neural decoder, embedders. |
ravana-v2/src/ravana_grace/ |
GRACE 20-phase cognitive governor (A–P). |
scripts/ |
Runnable entry points (chat, train, learn, benchmarks). |
experiments/ |
Research harnesses used by the benchmarks. |
tests/ |
pytest suite (ci / unit / integration / eval). |
docs/ |
Architecture, modules, training, benchmarks, development guide. |
The three src trees are one integrated system; scripts/ravana_chat.py
imports from all of them.
python -m pytest tests/ci/ -v --ci # fast critical-path job (~15 min, soak tests)
python -m pytest tests/unit/ -q # module-level
python -m pytest tests/integration/ -q # cross-module
python -m pytest tests/ --tb=short # full suitescripts/acceptance_ledger.py grades each cognitive module GREEN/RED with
quantitative metrics (identity strength, stance count, fact count, graph size,
save/load roundtrip, determinism checksum, strategy diversity, episodic buffer,
neuromodulator levels). Run it from the repo root:
python scripts/acceptance_ledger.pyExit code 0 = all GREEN, 1 = any RED. See the determinism contract above for the reproducibility guarantee.
CI (.github/workflows/ci.yml) runs pip install -e .[full,dev] then the
ci / unit / integration jobs on Python 3.10.
Results are current as of the latest commits. Historical values from prior experiments are archived separately.
| Metric | Result |
|---|---|
| Cross-domain transfer Top-1 | 100% (n=6) |
| Held-out Science Top-1 (post-sleep) | 93.8% (n=16) |
| Held-out Social Top-1 (post-sleep) | 80.0% (n=20) |
| Held-out vs baseline (Science) | 12.5% → 93.8% |
| Held-out vs baseline (Social) | 5.0% → 80.0% |
| Transfer probes Top-1/Top-10 | 59.5% / 73.8% |
| Sleep cycle conceptual accuracy | 90.2% |
| Graph size | find_similar p50 |
find_similar p95 |
|---|---|---|
| 1K nodes | 0.021 ms | 0.025 ms |
| 5K nodes | 0.043 ms | 0.051 ms |
| 10K nodes | 0.071 ms | 0.191 ms |
| Metric | Pre-ARC | Post-ARC |
|---|---|---|
| Composite quality | 0.395 | 0.394 |
| HONEST-abstinence | 0.600 | 0.700 |
| Confabulation rate | 0.000 | 0.000 |
| Salad rate | 0.000 | 0.000 |
| Verdict | ARC MAINTAINS QUALITY, IMPROVES HONESTY |
Results from a one-off measurement; see
docs/BENCHMARKS.mdfor how to reproduce.
| Metric | Result |
|---|---|
| Cross-domain transfer Top-1 | 75.0% |
| Cross-domain transfer Top-10 | 100% |
| Held-out Science Top-1 / Top-10 | 8.3% / 25.0% (n=12) |
| Within-domain triple top-10 | 80.9% |
| Lifelong forgetting (permuted MNIST) | 0% (with sleep) |
| Graph Inference P95 / P99 | 2.7 ms / 2.9 ms |
| Graph Peak Memory / Throughput | 0.3 MB / 556 QPS |
| W_rel Causal / Semantic Alignment | 0.68 / 0.55 |
See
docs/BENCHMARKS.mdfor how to reproduce each result. Seeexperiment_results/for raw JSON output files.
See docs/:
- Architecture — turn-level data flow and the three packages.
- Modules — what lives in each package.
- Training —
train.pymodes and the LingGen promotion gate. - Benchmarks — every benchmark/diagnostic script and what it measures.
- Development — layout, path shims, test commands, conventions.
- Entity-Location Recall — capturing + surfacing a named thing's whereabouts.
- Quantity Memory — capturing counts you disclose, answering "how many", totalling "in total", correcting online.
- Reverse Pet Lookup by Name — answering "who is wren to me?" by reverse-indexing the pet store by the name value.
- Agent Self-Stance — RAVANA forms, records, and recalls its own stance on a discussed topic (grounded in your view, attenuated, persisted), and stays honestly silent otherwise.
- Acceptance Ledger — per-module test-coverage grade (GREEN/YELLOW/RED) with real numbers.
- Honest Reporting Standard — four enforceable rules for every RAVANA claim.
RAVANA is evaluated end-to-end with scripts/evaluate_ravana.py, which trains a
fresh dim=64 engine on TinyShakespeare (25 passes, no live web) and runs all
nine benchmark batteries — each in an isolated engine so no benchmark leaks
facts into another — then evaluates the trained model against a same-scale nanoGPT
baseline across both architectural efficiency and task performance / cognitive capabilities.
Full per-case output and comparative metrics are exported to data/eval_results.json.
| Metric | nanoGPT (Transformer) | RAVANA (Cognitive Architecture) | Advantage |
|---|---|---|---|
| Parameters | 10,700,000 | 5,070,789 | 47.4% of nanoGPT params |
| Data Size (Tiny Shakespeare) | 1,115,394 chars | 1,115,394 chars | Same evaluation corpus |
| Param / Data Ratio | 9.60 params/char | 4.55 params/char | 2.1× higher data efficiency |
| Architecture | Causal Self-Attention (6L, 6H, d=384) | Neural Decoder (GRU+Attn) + Concept Graph | Hybrid neuro-symbolic substrate |
| Training Algorithm | Backpropagation (AdamW) | Local Predictive Hebbian Learning | Biologically plausible, forward-only |
| Optimizer Memory Overhead | ~171.2 MB (FP32 1st/2nd moments + grads) | 0.0 MB (in-place local updates) | Zero gradient/moment buffer overhead |
| Computation Graph Retain | Required for backward autograd pass | None (forward-only streaming) | Minimal memory footprint during learning |
| Metric | nanoGPT (Transformer) | RAVANA (Neural Decoder) | Notes & Analysis |
|---|---|---|---|
| Tokenization Scale | Character-level (65 chars) | Word-level (50,000 words GloVe) | RAVANA maps directly to semantic word embeddings |
| Next-Token Top-1 Accuracy | ~60.0% (over 65 characters) | 34.8% (over 50,000 words) | nanoGPT excels at local character sequence completion |
| Cross-Entropy Loss | ~1.47 nats/char | 2.32 nats/word | nanoGPT optimizes raw surface text log-likelihood |
| Perplexity | ~4.35 (per-char) | 10.18 (per-word) | nanoGPT yields fluent surface Shakespearean prose |
| Catastrophic Forgetting | >70% loss on sequential shift | <5% loss (with sleep replay) | Local Hebbian updates + sleep consolidation prevent overwrite |
| Inference Graph Latency | Dense Attention |
Sparse Graph Retrieval (P95: 2.7 ms) | Constant-time sub-graph walks vs quadratic attention context |
Both systems evaluated across the 9 cognitive evaluation batteries under isolated conditions:
| Benchmark / Capability Battery | nanoGPT | RAVANA | Δ Advantage | Architectural Rationale |
|---|---|---|---|---|
| Lamp Test (3-Premise Causal Reasoning) | 0.00 | 1.00 | +1.00 (RAVANA) | nanoGPT lacks causal graph logic; RAVANA uses unit propagation over premise graph |
| Self-Evaluation (Metacognitive Honesty) | 0.20 | 0.82 | +0.62 (RAVANA) | nanoGPT sycophantically confabulates; RAVANA abstains when epistemic confidence is low |
| Memory Consistency (MemFail) | 0.28 | 0.70 | +0.42 (RAVANA) | nanoGPT recency-overwrites; RAVANA maintains isolated slot bindings and truth values |
| Adversarial Robustness (AdvBench) | 0.08 | 0.40 | +0.32 (RAVANA) | nanoGPT compliantly auto-completes harms; RAVANA redirects via constructive means-end paths |
| Temporal Reasoning (TimeDial Cloze) | 0.31 | 0.55 | +0.24 (RAVANA) | nanoGPT uses surface n-gram cues; RAVANA bounds duration and timeline order in graph |
| Long-Term Memory (LoCoMo 10-Conv) | 0.14 | 0.34 | +0.20 (RAVANA) | nanoGPT context window truncates history; RAVANA stores long-term hippocampal episodes |
| Cross-Session Memory (LongMemEval) | 0.15 | 0.34 | +0.19 (RAVANA) | nanoGPT confabulates on unstated facts; RAVANA preserves cross-session entity slots |
| Logical Reasoning (LogiQA 4-Way MCQ) | 0.25 | 0.37 | +0.12 (RAVANA) | nanoGPT matches random chance (1/4); RAVANA uses HPC→PFC entailment resolution |
| Practical Consultation (Consult Advice) | 0.65 | 0.57 | -0.08 (nanoGPT) | nanoGPT's surface fluency produces coherent advice; RAVANA uses goal-directed means-end graphs |
| Overall Benchmark Average | 0.23 | 0.57 | +0.34 (RAVANA) | RAVANA outperforms by +0.34 overall (+148% relative capability improvement) |
Key Architectural Takeaways:
- Surface Fluency vs Grounded Cognition: nanoGPT excels at local surface perplexity (~1.47 nats/char) and next-character prediction because dense causal self-attention is an optimal statistical compressor of syntax. However, without an explicit episodic memory buffer or causal reasoning graph, it suffers from catastrophic forgetting (>70%) and fails at multi-step causal deduction (0.00).
- Parameter & Compute Efficiency: RAVANA achieves a +0.34 overall benchmark advantage with 47.4% fewer parameters (5.07M vs 10.7M) and zero optimizer memory overhead (~171.2 MB saved) by replacing global backpropagation with local Hebbian predictive coding and a typed concept graph.
- Fail-Closed Abstention: RAVANA monitors epistemic boundaries, choosing honest silence over confabulation when confidence is low — a core biological design principle that purely autoregressive transformers lack.
In addition to macro cognitive benchmarks, the isolated neuro-symbolic substrate (RLMv2) is evaluated against in-process feedforward and causal transformer baselines on discriminative relational tasks (scripts/benchmark_vs_transformers.py):
| Model | Architecture | Training Algorithm | Exact Params | Optimizer Overhead | Computation Graph Mode |
|---|---|---|---|---|---|
| RLMv2 (RAVANA) | Neuro-Symbolic | Local Predictive Hebbian | 156,870 | 0.00 MB | Forward-only (streaming) |
| Linear Baseline | Feedforward | Backpropagation (Adam) | 5,472 | 0.04 MB | Autograd graph retained |
| nanoGPT (Causal Transformer) | Causal Transformer | Backpropagation (Adam) | 12,096 | 0.09 MB | Autograd graph retained |
| MLP Baseline (2-layer) | Feedforward | Backpropagation (Adam) | 15,816 | 0.12 MB | Autograd graph retained |
-
Controlled Multi-Seed Ontology Ablation (5 seeds): Seed ontological priors confer a
$+4.0% \pm 8.0%$ generalization advantage over unseeded baselines under identical initializations and local updates. -
Continual Learning & Catastrophic Forgetting (
$A \to B \to C$ ): RAVANA achieves a 27.5% average retention loss via sleep consolidation and local Hebbian plasticity, retaining stable memory across sequential domain shifts. - Cross-Domain Transfer & Grounding: In the absence of a grounded semantic manifold (e.g. GloVe or ConceptNet), orthogonal token embeddings yield zero cross-domain transfer (0.0%), confirming that analogical projection fundamentally requires grounded semantic geometry rather than ungrounded token lookups.
For the complete experimental report, methodology, and raw JSON data, see reports/benchmark_transformer_comparison.md and docs/BENCHMARKS.md.
RAVANA guarantees bit-exact reproducibility under controlled conditions:
- Same seed + same inputs = same SHA256. Two
CognitiveChatEngineinstances with the samedim,seed, andbaby_modeparameters, fed the same sequence ofprocess_turn()calls, produce identicalstate_checksumvalues aftersave(). Verified byscripts/acceptance_ledger.py(determinism grade). - Same seed = same SHA256 is the contract. Different seeds produce different hashes — this is expected and correct.
- Offline mode (
RAVANA_OFFLINE=1) is required for reproducibility. Online mode fetches web content that varies between runs. - What is deterministic: identity state, stances, facts, beliefs, turn count, learning count, free energy, prediction error, neuromodulator levels, graph structure (nodes + edges), and the RNG state.
- What is NOT deterministic across processes: the
ConceptGraphobject's in-memory representation (it is serialized to a string on save and rebuilt on load). Thestate_checksumexcludes the graph for this reason — seeengine_persistence.py:_checksum_state. - Checksum scope: the
state_checksumis a SHA-256 fingerprint of the durable cognitive state (identity, user model, beliefs, turn count, etc.), excluding the graph and the checksum itself. It is stored in the pickle and in a.shasidecar file.
To verify reproducibility:
import pickle
# Run 1: create engine, process turns, save
eng1 = CognitiveChatEngine(dim=64, seed=42, baby_mode=True)
eng1.process_turn("i love coffee")
eng1.save()
# Run 2: same seed, same turns
eng2 = CognitiveChatEngine(dim=64, seed=42, baby_mode=True)
eng2.process_turn("i love coffee")
eng2.save()
# Compare checksums
with open(eng1._save_path, "rb") as f:
s1 = pickle.load(f)
with open(eng2._save_path, "rb") as f:
s2 = pickle.load(f)
assert s1["state_checksum"] == s2["state_checksum"] # passes- Fail-closed grounding — abstain rather than confabulate.
- No fixed thresholds where a distribution exists — gates are data-derived.
- Continuous, curiosity-driven learning — what to learn is selected from prediction error, novelty, and contradiction, not a fixed list.
- Learning without backprop —
ravana_mlminimizes free energy and consolidates duringsleep_cycle().
RAVANA is licensed under the Oxiverse Community License (OCL).
The full, always-current license text — including the current version, the commercial-licensing terms (closed-source self-hosting and open-source self-hosting paths), and the privacy-by-design requirements — is published at:
https://oxiverse.com/license
That page is the single source of truth (it renders the canonical LICENSE directly), so read it there rather than relying on a copied snapshot.
Quick summary:
- Non-Commercial community use (student projects, portfolios, research, internal evaluation) is permitted.
- Modifications to the core codebase must remain under OCL.
- Commercial use — SaaS, paid products, hosted offerings, or closed-source self-hosting — requires a separate Commercial License from Oxiverse (inquire at likhith@oxiverse.com).
- Independently developed apps that merely consume the hosted API/platform are NOT derivative works and may be licensed under your choice of OSI-approved license or OCL.
- Privacy-by-Design is a non-negotiable baseline for all use.
This repository does not embed its own copy of the license to avoid drift.