Type classification for lemmas

Every lemma carries exactly one type. The type vocabulary has two sorts: ontological types, for lemmas that encode a concept, and procedural types, for lemmas that encode an instruction for processing. Which sort a lemma takes is the first decision the annotator makes; everything else follows from it.

A lemma here is a lexeme: a citation form paired with a part of speech, not a bare string. A lexical unit (LU) is a lemma paired with a frame, carrying a sense description — the relation is N:M, exactly as in Berkeley FrameNet. The type belongs to the lemma and only to the lemma. The LU inherits it and carries no type of its own, so creating an LU involves no typing decision at all.

What this layer claims, and what it does not#

Four scoping commitments govern every assignment in this document.

  1. The type is a property of the lemma — a lexeme, text plus POS — not of a word form or a referent. POS already does the construal-fixing work: colar/VERB and colar/NOUN are different lexemes with different meanings before any frame is consulted.
  2. The type classifies a conceptual profile, not the furniture of the world. No claim is made about what exists. This follows the grounding ontology: DOLCE is a Descriptive Ontology, an inventory of particulars as conceptualized by natural language, explicitly neutral on metaphysics. Typing a lemma against DUL types a conceptualization against a catalogue of conceptualizations.
  3. The type is a coarse, defeasible default, serving as the input to constructional composition and coercion. A construction may override it — override is the interesting data, not a counterexample, since a function needs an argument. The default lives in the lexicon and is defeated in use.
  4. The type is a comparative concept, in Haspelmath's sense: an FNBr-internal instrument for cross-frame, cross-resource and cross-linguistic alignment — not a descriptive category of Portuguese, and not a metaphysical claim.

Because the type is defeasible (commitment 3), the class-compatibility check in §8 flags a marked combination worth review, not an error.

1. Parts of speech and the locus of the type#

1.1 Why retain POS at all#

The theoretical case against parts of speech as a semantic category is strong: syntactic categories are language-particular and construction-derived (Croft, Radical Construction Grammar), and do not translate across languages as descriptive categories (Haspelmath 2010). This is why the type layer exists and why it is not a POS layer.

POS is nonetheless retained: it is load-bearing for lemmatization and morphological analysis, MWE and terminology identification, corpus query (word sketches are defined over POS patterns), and interoperability with UD, LMF and OntoLex-Lemon/LexInfo.

The architecture follows OntoLex-Lemon: POS is morphosyntactic metadata on the lexical entry; the type is the conceptual layer, linked by semantics-by-reference. Both sit on the lemma; neither pretends to be the other.

1.2 POS is part of lemma identity#

POS is a property of the lexical entry = (lemma string + POS), and the lemma key must include it:

lemma(id, text, pos, lang, homograph_idx, type_id)
        UNIQUE(text, pos, lang, homograph_idx)
form (id, text, lemma_id, feats)      -- the paradigm of ONE lemma entry
lu   (id, lemma_id, frame_id, sense_description)

Keying lemma on text alone conflates distinct lexemes. Colar is two lemmas: colar/VERB ("to glue"), whose paradigm contains colo, colei, colando; and colar/NOUN ("necklace"), whose paradigm is colar, colares. With a text-only key, a parse of Eu colo o papel would resolve colo to a single colar row and make the noun sense reachable from a verb form — a lemma that does not exist.

Two consequences:

  • Form lookup is set-valued. form.text → {lemma_id} returns candidates. POS is what filters them, but POS lives on the token, not on the form string: parsing the sentence assigns the token a UPOS, and matching that against the POS carried by each candidate lemma picks the right one. Colo is itself ambiguous off any list — it is a form of colar/VERB and also a noun in its own right (colo/NOUN, "lap") — so the string alone does not resolve it; the token's POS from the parse does, no semantic reasoning needed.
  • When UPOS is absent or unreliable, do not guess at the lemma layer. Return the candidate set and let frame matching decide. POS narrows; it does not adjudicate.

homograph_idx distinguishes lemmas that share text and POS: genuine homonymy (manga fruit/sleeve) and type divergence, where one form under one POS carries readings of different types. Default 0. It is not used for polysemy within a single type — that is what LUs are for (§1.4).

For MWEs, pos is X and entry_type ∈ {word, mwe, affix} carries the word/MWE/affix distinction. UD's own X tag is defined for "words that for some reason cannot be assigned a real part-of-speech category" — exactly the case for dar um jeito. It asserts that no single POS applies, rather than asserting a wrong one. This is also where the type layer outperforms POS: an MWE has no coherent part of speech but a perfectly coherent ontological type.

X, not null. A null POS propagates: view_lu derives an LU's display name as concat(lemma.name, udpos.suffix), and concat() returns null when any argument is — so a null POS silently blanks the rendered name. X carries a real suffix, so dar um jeito renders as dar um jeito.x. Every lemma carries a POS; a lemma arriving without one is assigned X rather than left null. (lemma.idUDPOS remains nullable in the schema for legacy reasons, but nothing relies on that.)

1.3 Why the type belongs to the lemma#

Three reasons, one theoretical and two practical.

The lexeme is already a form–meaning pairing. Once POS is part of lemma identity, most of the construal-fixing work is already done before any frame is consulted. Colarᵥ and colarₙ are separate lemmas with separate types; so are dependenteₙ (role) and dependenteₐ (value). Putting the type on the lemma–frame pairing instead would record, on every LU, a fact constant across all of a lemma's LUs.

The type is available before frame assignment. A parser can type a token as soon as it resolves the lemma, with no frame in hand — which is precisely when a type is most useful as a filter.

It matches the FrameNet data model. The LU table is a lemma × frame association with a sense description. A mandatory type column there would make the LU a second lexical entry rather than a pairing.

1.4 One lemma per (text, POS, type)#

Where one form under one POS carries readings of different types, each type gets its own lemma row, distinguished by homograph_idx. Readings that share a type stay in one lemma and are distinguished as LUs.

One lemma per (text, POS, type).

This is the only mechanism for type divergence in the specification — there is no LU-level type and no override.

  • Creating an LU involves no typing decision. The annotator pairs a lemma with a frame and writes a sense description; the type arrived with the lemma.
  • lemma.type_id is correct by construction. It is non-null, single-valued and never wrong for any LU beneath it, because a divergent reading is by definition a different lemma.

The split is bounded by type, not by sense count: de has at least six distinguishable readings but only two types, so it yields two lemmas, not six (§6.2).

Three case shapes are known:

  • Cross-sort polysemy. Negação names the act of denying (event, in an @agentive frame) and the propositional operator (polarity) → negação₀ event, negação₁ polarity.
  • Product nominalization. Pintura as the activity is event; as the canvas on the wall it is object. Same for construção, tradução, invenção, criação, publicação — deverbal nominals whose event produces an artifact that outlives it. Not all deverbal nominals: see §4.
  • Adpositions. De, a, em, por, com, para, sobre — the largest group, worked through in §6.2.

Not for:

  • Construal shifts. A lemma can be referred to, or otherwise used, in a way that leaves part of its normal profile inactive without the stored type changing — the type is a fact about the lemma; the construal is a fact about the occurrence. Pai used to pick out an individual with no relation in play is still the role lemma: the relation is simply not invoked by that use. Löbner's determiner-triggered readings are the best-studied case, but the mechanism is not limited to determiners — see the next bullet. Per commitment 3, a construal shift is defeasance in use, not a second type — and, since nothing about the lemma itself changed, not a second lemma either.
  • Substantivized adjectives used elliptically. O bonito, um americano pick out an entity the adjective characterizes — a person, a car, whatever the covert head recovers as — via the construction that lets any adjective head an NP under a determiner with its head noun elided. This is a construal shift like the one above: bonito and americano stay value, and which entity the value is predicated of is a frame-level fact (§7), not a lemma-level one. The test: can the use be regenerated as "[some contextually recoverable sortal] that is ADJ," with the adjective keeping its ordinary value? If the covert head is freely recoverable and swappable by context, it is a construal shift and the adjective lemma is untouched. If the "noun" reading is fixed regardless of context and carries content the adjective does not — typically a specific institutional or relational profile — the use has lexicalized, and §1.2's POS split applies instead: dependenteₙ ("meu dependente", a role — someone for whom a specific provider must exist) is a different lemma from dependenteₐ ("ele é dependente", value), not a construal shift of it (§7).
  • Polysemy within a type. Genitive, material and partitive de are one predication lemma with several LUs. Two frames do not imply two lemmas; two types do.

Splitting is the same operation as homonymy handling, deliberately: manga fruit/sleeve and de₀/de₁ produce identical rows. Annotators need not decide whether two readings are "the same word" — once the type is settled, that judgement has no consequence for the data.

A split records no relation between the two lemmas. Where relatedness matters — negação act/operator, pintura activity/canvas — it belongs at the lexical layer as an explicit link (OntoLex-Lemon's vartrans module), not as a shared row.

1.5 POS and type are not orthogonal#

Deverbal nominalizations and participles show strong type/POS correlation, because category is assigned in a structure that also determines event structure. The claim is narrower: the type is not derivable from POS and not reducible to it. How much independent information it carries is answerable by a POS × type contingency table, since both columns live on the same row.

2. The two sorts: conceptual and procedural encoding#

Not every word encodes a concept. Quem, não, portanto, o have meanings, but those meanings are not things classifiable as objects, events or qualities — the "world thing" is computed in context, if at all.

Relevance Theory supplies the distinction. Blakemore (Semantic Constraints on Relevance, 1987) and Wilson & Sperber ("Linguistic form and relevance", Lingua 1993) distinguish conceptual encoding — the expression contributes a constituent to the proposition expressed — from procedural encoding — the expression encodes an instruction for processing: a constraint on inference, reference resolution, or how to build or access a representation. Discourse connectives, indexicals, polarity, and most grammatical morphemes are procedural on this analysis.

FNBr adopts the distinction as the top branch of the type vocabulary:

Sort Encodes Gets an ontological type Gets a DUL reference
conceptual a concept yes (§3) yes
procedural an instruction for processing no no

Procedural lemmas carry no DUL reference — there is nothing in the conceptualization inventory for them to reference, because they name an operation on a conceptualized particular, not the particular itself.

3. The ontological types#

Six types. Each grounds in a DUL class. There are no parameters.

Type DUL grounding Criterion Examples
event dul:Event occurs or happens: it exists in time, has (or can have) temporal parts, and does not persist unchanged destruição, aquecer, correr, quebrar, derreter, existir
object dul:Object a sortal — it supplies a criterion of identity — whose instances need no other entity to exist cadeira, água, pessoa, documento, lei, ideia, criança, adulto
role dul:Role ⊑ dul:Concept a sortal whose instances are externally dependent: some other entity, not a part of them, must exist comprador, jogador, cliente, filho, pai, proprietário
relation dul:Relation ⊑ dul:Description a reified relation type — a Description that defines the argument positions (roles) it relates parentesco, paternidade, posse, casamento, adjacência, anterioridade
quality dul:Quality a dimension — a gradable or categorial axis — that inheres in some bearer cor, tamanho, beleza, inteligência, temperatura, cansaço, postura
value dul:Region ⊑ dul:Abstract a region or point within a quality's dimension, nameable without presupposing any bearer azul, quente, bonito, inteligente, cansado, quebrado, sentado, belamente

The criteria draw on rigidity, sortality, external dependence, and the dimension/region distinction — defined in full in Appendix §A.1.

The reference principle#

The lemma references a conceptualization, never a particular and never a formal object:

  • comprador references dul:Role — a Concept that classifies persons — not any person.
  • parentesco references dul:Relation, a Description — the relation type — not any particular kinship instance.
  • adjacência and anterioridade reference the commonsense dul:Relation, not the formal counterpart under dul:FormalEntity.

This is what "semantics by reference" means for a lexicon: the lemma points at a conceptualization, which is exactly what a word does.

role and relation are two halves of one structure#

dul:Relation is a Description; the Concepts filling its argument positions are dul:isDefinedIn that Description:

FNBr object DUL DnS relation
relational frame dul:Relation ⊑ Description —
FE of that frame dul:Role / dul:Concept dul:isDefinedIn → the Description
relation lemma (parentesco) references the Relation —
role lemma (filho, pai) references a Role isDefinedIn the same Relation

Parentesco names the description; filho and pai name Concepts it defines — not two unrelated primitives that happen to sound alike.

The instance layer, where needed, is dul:Situation: the particular fatherhood holding between two individuals in an annotated sentence is a Situation that dul:satisfies the Relation. Lemma layer → Description; annotation layer → Situation — the same coarse-identity/fine-situation asymmetry the whole model runs on (§7).

Converse pairs fall out of the structure. Both role and relation are grounded in DnS, so both pai and filho are isDefinedIn the same parentesco Relation, and the converse pairing is readable off that link.

Rigidity is not used as a typing axis, and abstract nouns do not get a type of their own — both are informational notes rather than operative rules; see Appendix §A.2 and §A.3.

4. Conceptual diagnostics#

The first question: conceptual or procedural?#

Does the lemma contribute a constituent to the proposition, or instruct the hearer how to process one? If the latter, go to §5 and stop.

object / role#

Both are thing-like; one binary test separates them: external dependence.

Type Examples
not externally dependent object pessoa, cadeira, documento, criança, adulto
externally dependent role comprador, filho, presidente, proprietário

Is there an entity y, not a part or constituent of x, whose existence is required for x to be P?

Comprador → yes (a purchase, a seller). Filho → yes (a parent). Documento → no. Criança → no: being a child requires no second entity, so it is object, not role, even though an individual stops being one (Appendix §A.2).

Mere participation is not dependence. Documents get signed, filed and sent, but no particular event must exist for a document to be a document. If participation sufficed, every entity noun in the lexicon would be a role.

Two corpus diagnostics, for cases where the ontological test is hard to apply — both operationalize Löbner's [R] (Appendix §A.1):

  1. Bare predicative: ?Ela é irmã vs Ela é médica. A +R concept resists the bare predicative without a recoverable relatum.
  2. De-genitive reading: does de X give an argument (pai de João) or a modifier (carro de João)? Argument readings indicate +R.

These are annotatable and scorable, so the role boundary can be validated rather than defended by introspection.

event / object in deverbal nominals#

A deverbal noun is an event by default and stays one. Nominalization is a constructional fact: it lets an event concept be referred to, counted, determined and modified like a thing, without changing what the concept is. Duas destruições, a destruição foi total, aquela destruição are all the event under nominal packaging. Countability, determination, pluralization and adjectival modification are never evidence for object.

The one case that diverges is a lemma with a product sense: a participant of the event — typically its Result or Product — with existence independent of the event, persisting after it ends.

Does the lemma have a sense denoting a thing that the event produced, rather than the event itself under some construal?

Three tests, all passed by the product sense and failed by the event sense:

  1. Spatial location. A pintura está na parede. / *A destruição está na parede.
  2. Physical or material predicates. A tradução tem 300 páginas. / *A destruição tem 300 páginas.
  3. Persistence. The product exists once the event is over; the event does not.

Pintura, construção, tradução, invenção, criação, publicação pass. Destruição, aquecimento, chegada, corrida, viagem fail: they have no product, and their apparent "result" readings are the event's Result or Effect FE reached metonymically — a frame-layer fact, not a type-layer one.

Grimshaw's complex-event/simple-event/result trichotomy is not used here: it separates nominals by argument structure, not by type (Appendix §A.4).

quality / value#

Ask whether the lemma names the dimension or a region within it (Appendix §A.1). Altura/alto, temperatura/quente, beleza/bonito, cansaço/cansado are each a dimension/value pair, decided the same way every time.

No entity is presupposed by either. A temperature scale exists, and a point on it is nameable, whether or not any object is currently at that temperature — está quente hoje, said of the weather, names no entity at all. What links a value to the entity bearing it is the job of the Attributes schema (its Entity, Value and Attribute elements), not of the lemma's type.

The type does not encode causal origin. Whether a quality resulted from an event (cansaço, force-maintained postura) or not (beleza, inteligência) plays no role in the classification: both are quality, and their values — cansado, quebrado, sentado on one side, bonito, inteligente on the other — are both value. That distinction belongs to the situation, which the frame layer already records (a STATIVE schema vs an Attributes schema plus a relation to the producing event); it is the same asymmetry already accepted for destruição, which stays event while its frames range over @agentive, @change and @stative without the lemma splitting. See Appendix §A.5 for the full argument, including why no reliable test separates a "resultant" quality from a non-resultant one.

Referential use does not create a new type. O bonito, um americano refer to an entity characterized by the value without the lemma becoming a different kind of thing — see §1.4, "Substantivized adjectives used elliptically," for the construal-shift diagnostic that tells this apart from genuine lexicalization (dependenteₙ/dependenteₐ, §7).

relation / role#

The tie is relation (paternidade, posse); the argument position named as such is role (pai, proprietário). Arity and converse evidence (pai-de ↔ filho-de, posse ↔ pertence-a) confirm a relation exists; the link between the two is dul:isDefinedIn (§3), not a parameter.

A free-standing referential use of a role noun — pai picking out the individual who walked in, with no relation in play — is a construal shift, not a second type. The lemma remains role (§1.4).

5. Procedural types#

Nine types. Procedural lemmas take no DUL reference and are exempt from the class-compatibility check (§8).

Type Encodes Examples
reference how to resolve a referent — deictic, anaphoric, definiteness o, este, ele, quem, aqui, hoje, agora
quantification how to quantify over a domain todo, cada, algum, vários, ambos, sempre
polarity assert or deny instantiation sim, não, nunca, jamais, nada, nenhum, sem
degree how to locate a value relative to a standard muito, pouco, mais, menos, bastante, demais, tão
modality epistemic or deontic status of the proposition, with no scale named talvez; modal auxiliaries (poder, dever, ter que); free mood markers
connection inferential or discourse relation between units contudo, portanto, porém, entretanto, pois, se
focus information-structural instruction só, também, até, mesmo, inclusive
predication how to build the predication; contributes no content copulas (ser, estar), auxiliaries, support verbs, case-marking de/a/por
interaction how to manage the exchange — appeal, stance display, information source, hedging né, hein, viu, tá, ué, nossa, olha, tipo, assim, diz que

The nine types are organized by functional category; their theoretical provenance and relation to Relevance Theory is in Appendix §A.6.

Notes on the boundaries#

degree vs value. The criterion is whether the lemma names a region or instructs where to locate one relative to a standard. Quente, morno, azul name regions → value. Muito, pouco, bastante locate a region against a contextually supplied standard → degree.

Prepositions are not a single case. ADP splits across both sorts: contentful spatial and temporal adpositions are relation, governed case-markers are predication, and some are polarity or connection (§6.2).

modality vs value: does the lemma name a scale? Epistemic adverbs with a lexicalized dimension are value, not modality. Provavelmente, certamente, possivelmente and seguramente each name a region on the probability/certainty dimension (probabilidade, certeza), have a gradable adjectival counterpart (provável, certo), and stand to it exactly as belamente stands to beleza (§6.1). Three tests separate them from modality:

Test value — provavelmente modality — talvez
degree modification muito provavelmente ✓ *muito talvez ✗
lexicalized dimension and adjective probabilidade, provável ✓ none ✗
embeds under negation não é provável que… ✓ *não talvez… ✗
governs mood no yes — subjunctive

The first test is decisive on its own: degree modifies value by definition, so anything muito can stack on is a value. Talvez fails all three and governs the subjunctive — an instruction to the grammar rather than a region on a scale. Modal auxiliaries form a third group with no scale at all and stay in modality.

The copula and support verbs are predication. They instruct how to assemble a predication and contribute no concept of their own.

interaction is about the exchange, not the proposition. It covers interjections (ué, nossa, ai), interactional and tag particles (né, hein, viu, tá), attention markers (olha, escuta), hedges and approximators (tipo, assim), and quotative or evidential markers (diz que, dizem que) — whatever manages the interaction rather than contributing to or operating on the proposition expressed.

Three boundaries:

  • vs connection. A connective relates two propositions inferentially (portanto, contudo); an interactional particle relates the utterance to the hearer (né seeks confirmation, viu secures uptake). Test: does removing it leave an inferential gap between two units, or only a blunter utterance?
  • vs modality. modality marks the speaker's own commitment to the proposition (talvez, poder); evidentials mark where the information came from (diz que). FNBr keeps evidentials in interaction because they invoke the hearer's assessment of the source rather than stating the speaker's confidence (Appendix §A.7).
  • vs focus. Só, até, mesmo operate on the proposition's information structure and are focus. Tipo and assim as hedges signal approximation in the exchange rather than restructuring the proposition.

interaction lemmas evoke pragmatic frames like any other lemma; they are exempt from the class-compatibility check of §8 along with the other procedural types.

6. Parts of speech the type layer cuts across#

Some UPOS tags collect lemmas with almost nothing semantically in common. The type layer is expected to cut across them; ADV and ADP need worked guidance.

6.1 Adverbs#

"Adverb" is a part-of-speech wastebasket, not a semantic class — it collects manner, place, time, frequency, degree, epistemic and connective words with almost nothing in common. Under the two-sort vocabulary, the class splits cleanly, with most members turning out to be procedural.

Adverb class Encodes Type
Manner (belamente, rapidamente) a region on a quality dimension, predicated of an event value
Time extent (recentemente) a region on a temporal-distance dimension value
Frequency, graded (raramente, frequentemente) a region on the frequency dimension value
Degree / intensity (muito, pouco, bastante) locate a value against a standard degree
Universal frequency (sempre) quantification over times quantification
Epistemic, scalar (provavelmente, certamente, possivelmente) a region on the probability/certainty dimension value
Epistemic, non-scalar (talvez) epistemic status, no scale named; governs subjunctive modality
Epistemic stance as a concept (dúvida, incerteza, certeza) the dimension itself quality
Place and deictic time (aqui, hoje, agora) resolution instruction reference
Conjunctive (contudo, portanto) discourse relation connection
Affirmation / negation (sim, não, nunca, nada, nenhum) polarity polarity

Manner adverbs are values, not dimensions. Belamente stands to beleza exactly as bonito does — a region on the same dimension, predicated of an event rather than an entity. What the value is predicated of is a frame-level fact, not a type-level one.

Deictics are procedural, not relations. Aqui and hoje encode an instruction to resolve a referent against the utterance situation — the textbook case of procedural encoding. Their image-schematic character is carried by what their frames ground in (FIGURE_GROUND, DEICTIC).

Zero is polarity, not a low value. Nada and nenhum look like values on the quantity dimension, parallel to pouco. But a low value presupposes existence — pouco asserts some quantity, just a small one — while nada and nenhum deny existence outright (¬∃x). The same line separates raramente (a low but positive frequency, still asserting the event occurred) from nunca (a negative existential over time points). A zero reading on any dimension negates that dimension's existential presupposition rather than measuring a small amount of it. Nada is polarity.

The polarity relation is not the act. Não as propositional polarity is polarity; the act of affirming or denying (afirmar, negar) is an event in an @agentive frame. Negação carries both readings under one form and one POS, and is split into two lemmas accordingly (§1.4).

6.2 Prepositions and adpositions#

ADP is one UPOS covering lemmas on both sides of the conceptual/procedural line; the split does not follow the tag at all. A preposition may name a relation, mark a case, express polarity or carry a discourse connection.

Use Examples Type
Spatial or temporal configuration sobre, sob, entre, contra, durante, após, dentro de, ao lado de, antes de relation
Privative sem polarity
Causal, concessive, conditional por causa de, apesar de, devido a, em vez de connection
Governed by the predicate (case-marking) de in gostar de, a in obedecer a, em in acreditar em, com in contar com predication
Marking a role the frame already supplies por (passive agent), com (instrument), para (beneficiary) predication

Three diagnostics:

  1. Contrastive substitution. Holding the governing head constant, can another preposition replace it with a truth-conditional difference? O livro está sobre / sob a mesa — yes, so relation. *Gostar em — no, so predication.
  2. Questionability. Can the PP answer a wh-question with the preposition inside the answer? Onde? — Sobre a mesa. Governed prepositions fail — there is no question whose answer is de in gostar de.
  3. Government. Is the preposition listed in the valence of the head that selects it? If yes, predication regardless of what it means elsewhere.

Diagnostic 3 overrides the other two when they conflict: government is observable in the valence data, while the other two rest on judgement.

Governed prepositions are markers, not frame-evoking lemmas. A case-marking preposition contributes no concept and gets no frame — its place is the valence pattern, as the marker of an argument position (the role synsem:marker plays in OntoLex-Lemon). predication records that the lemma is real and type-bearing but conceptually empty; which head governs it is recorded at the valence layer, not by a frame link.

Applying the rule. Adpositions are the largest application of §1.4 — one lemma per (text, POS, type). Splitting stays bounded because readings that share a type do not split: de has at least six distinguishable readings but only two types, so it yields two lemmas, not six; genitive, material and partitive are LUs of the same predication lemma because they are all markers.

Lemma Type Covers Examples
de₀ predication government; genitive; material; partitive gostar de música; o carro de João; mesa de madeira; um copo de água
de₁ relation source and origin, spatial or temporal veio de São Paulo; de segunda a sexta
a₀ predication government; dative and recipient marking obedecer a uma ordem; dei o livro a João
a₁ relation goal, direction, temporal location fui a Roma; às três horas
em₀ predication government acreditar em algo; pensar em você
em₁ relation containment, location, temporal location na mesa; em 2020; no Brasil
por₀ predication passive agent; government; price marking destruído por João; optar por outro; comprei por dez reais
por₁ relation path, duration, cause passou por São Paulo; por três horas; preso por roubo
com₀ predication government; instrument and manner marking contar com você; cortou com a faca; com cuidado
com₁ relation comitative, accompaniment foi com Maria; café com leite
para₀ predication government; beneficiary and recipient marking olhar para a janela; comprei um livro para João
para₁ relation goal, direction fui para Roma; para o norte
para₂ connection purpose, linking two events saiu para comprar pão
sobre₀ relation superposition o livro está sobre a mesa
sobre₁ predication topic marking um livro sobre linguística
sem₀ polarity privative — no split, one type covers every reading sem açúcar; saiu sem pagar

Sobre splits even though it looks like a paradigm case of a contentful preposition — the split is not a property of "grammatical" adpositions only. Sem does not split at all — the split is driven by type divergence and nothing else.

Disambiguation. Splitting means UPOS filters nothing between de₀ and de₁ — they share both text and POS. Disambiguation comes from the syntactic context, using evidence already in the UD parse:

  1. Government. If the adposition's head is a predicate that selects it — the case dependent of an obl/obj whose head lists that preposition in its valence — choose the predication lemma. This decides most tokens.
  2. Complement semantics. A place or time complement favours the relation lemma; an infinitival complement under para forces para₂.
  3. Default. Where neither fires, prefer the predication lemma — the more frequent reading for de, a, em, por, com and para, and the less damaging error since it attributes no content that is not there.

Extended discussion of the accuracy trade-off this moves to the parser is in Appendix §A.8.

Nothing here is specific to adpositions — the same rule produces negação₀/negação₁ and pintura₀/pintura₁ in the content lexicon (§1.4). Adpositions are simply where it fires most often: a small set of very frequent forms each span the conceptual/procedural line. The split count is concentrated — a handful of function words account for most of it.

Complex prepositions are MWE lemmas. Dentro de, ao lado de, por causa de, apesar de take entry_type = mwe and pos = X (§1.2). Their type follows the table above and is often connection rather than relation, which is a further reason not to read the type off the tag.

The type cuts across ADP/SCONJ as well as within ADP. Por causa de (ADP) and porque (SCONJ) are both connection; para + infinitive sits on the boundary — the tag predicts the type in neither direction.

Consequence for §8. The structural check requires a relation lemma's frames to have FEs lexicalized by role lemmas. Spatial prepositions systematically fail this: the FEs of a locative frame are Figure and Ground, and Portuguese has no role lemmas naming them. The check is therefore scoped in §8 to relational frames whose FEs have role lexicalizations.

7. Why coarse: the three layers#

The asymmetry is the most important property of this design; with the type on the lemma it reads as three layers:

Layer Grain Carries
lemma coarse identity the type — what kind of thing the lexeme names
LU sense the lemma–frame pairing and a sense description — no type
frame situation the namespace, the participant structure, the grounding schema
  • Destruição is one lemma typed event. Its LUs may sit in an @agentive frame (destruction as cause), a @change frame (destruction as change of state) and a @stative frame (destruction as resulting condition) — without contradiction and without splitting the lemma.
  • Aquecer is one lemma typed event, with LUs in Causar_mudança_de_temperatura and Mudança_de_temperatura. Same type, different frames — the ordinary case.
  • value behaves the same way across @attribute, @stative and @entity. Americano names a region on the origem dimension, full stop, regardless of what it is predicated of: a person (um americano), a thing (um carro americano) or an event (o ataque americano). frm_people_by_origin is an @entity frame built over value lemmas, and it is not an exception — the live data has dozens more (frm_people_by_religion, frm_people_by_vocation, frm_wealthiness, frm_expertise, and around twenty-five others). This is a construal shift, not a split (§1.4): the lemma stays value and the referent varies with what the construction supplies.
  • POS differences split the lemma, not the LU. Dependenteₙ ("meu dependente") is role; dependenteₐ ("ele é dependente") is value — two lemmas by §1.2, each with its own type and its own LUs. Unlike americano above, this is not a construal shift: the noun names a fixed institutional role, not "person" as a default referent for the value (§1.4).
  • Type divergence splits the lemma too. Where one form under one POS carries two types, it is two lemmas by §1.4. Nothing in the LU table records a type, so no divergence can be hidden there.

Construal shifts change neither. Ele é pai shifts a functional concept to a sortal predicate without changing the underlying concept (Löbner); o bonito and um americano refer to an entity a value lemma characterizes without becoming sortals themselves. Neither is type ambiguity, neither spawns an LU, and neither splits the lemma — the declared type stands, per commitment 3. §1.4 has the general statement and the diagnostic that tells a construal shift apart from a genuine split like dependenteₙ/dependenteₐ above.

8. What this layer validates, and what it cannot#

All checks run on lemma.type_id. There is no second source to reconcile: an LU's type is its lemma's type, always.

The type licenses a class-compatibility check: does the frame's namespace admit this type at all?

Namespace Admits
@agentive, @change, @phenomenon, @process, @undergoing event
@experiential event
@stative event (the maintaining verbs — permanecer, manter — are eventive lemmas like any other verb evoking an eventive frame), quality, value
@attribute quality, value
@entity object, role, value
@relational relation, role
— procedural types are exempt

A second, structural check:

For every relation lemma, the frames its LUs evoke should have FEs lexicalized by role lemmas, and those roles should be isDefinedIn that same frame-as-Relation.

Scope. The check applies only to relational frames whose FEs actually have role lexicalizations. Spatial and temporal frames are excluded: their FEs are Figure and Ground, which no Portuguese lemma names as a role (§6.2).

A relation lemma whose frames have no role lexicalizations is either mistyped or sitting in a frame missing its participant LUs. A role lemma whose frames are not Relations is a typing error.

A third check, available only because the type sits on the lemma:

A lemma whose LUs span frames with incompatible admit-sets is a candidate for a split (§1.4).

What the checks do:

  • Flag cross-class category errors: a quality lemma in an eventive frame, an object lemma in a stative frame, a relation lemma in an agentive frame. Automatable, high-precision. Per commitment 3, a flag is a marked combination for review, not a verdict.
  • Silent on within-class distinctions: @agentive vs @change vs @phenomenon all admit event, so the check says nothing about which is right.
  • Silent on wrong-sense errors: only an annotator working against a real sentence can choose which frame an occurrence belongs to.
  • Does not assign types automatically. Which of the six ontological types, or whether a lemma is procedural at all, is annotator judgement guided by §4–§6.

Appendix#

Secondary material: full definitions, theoretical provenance, extended diagnostics, design notes, and open questions. None of it is required to classify a lemma, but §A.1 underlies the criteria in §3.

A.1 Definitions#

Six notions from formal ontology and lexical semantics, used by the diagnostics throughout this document and by the criteria in §3.

Rigidity (Guarino & Welty, OntoClean). A property is rigid if essential to all its instances: anything that is P must be P for as long as it exists (pessoa, cadeira). It is anti-rigid if every instance could cease to be P and survive (comprador, criança, cansado), and semi-rigid if essential for some instances and accidental for others. Anti-rigid properties cannot subsume rigid ones — why pessoa is not a kind of agente. Rigidity is defined here but is not a typing axis (§A.2).

Identity and sortality. A property supplies identity if it carries a criterion for saying "the same one again"; such properties are sortals, answering what something is. Pessoa supplies identity; comprador inherits it from pessoa; vermelho has none. Every object instantiates exactly one identity-supplying kind. Both object and role are sortals (§3); what separates them is external dependence, not sortality.

Unity. A property carries unity if all its instances are wholes under one unifying relation (functional, topological, morphological, intentional). Pessoa carries functional unity; água is anti-unity — an arbitrary sub-portion of water is still water. FNBr does not type on unity; it explains why mass nouns need no separate type (§A.3).

External dependence. A property is externally dependent (+D) if every instance requires the existence of some entity y that is not a part or constituent of it. Comprador requires a purchase and a seller; filho requires a parent; cadeira requires nothing. This is the diagnostic that separates object from role (§3, §4) and replaces constitutivity.

Relationality and uniqueness (Löbner, "Concept Types and Determination", Journal of Semantics 28(3), 2011). A noun concept is relational [+R] if it has an inherent argument slot for a relatum (irmã, filho), and unique [+U] if it determines at most one referent, absolutely or relative to its relatum (sol; pai relative to a child). Crossing the two yields sortal [−R−U], individual [−R+U], relational [+R−U] and functional [+R+U] concepts. FNBr does not type on [U], but both features supply corpus- observable diagnostics (§4) where the ontological test is hard to apply.

Dimension and region (Gärdenfors, Conceptual Spaces; The Geometry of Meaning). A quality is a dimension — an axis if gradable, a class if categorial. A region is a value the dimension can take. Temperatura is a dimension; quente is a region within it — the distinction underlying quality vs value (§3, §4).

A.2 Why rigidity is not a typing axis#

Rigidity (§A.1) is not used to type anything. Anti-rigid intrinsic sortals — criança, adulto, novato, defunto, concepts an individual can stop falling under while continuing to exist, with no second entity required — are typed object, the same as rigid sortals like pessoa.

The reason is annotator reliability. The dependence test (§4) asks a single existence question: is there some entity, not a part of this one, that must exist? Rigidity asks a modal question: could this individual cease to fall under this concept and still exist? Modal judgements of that kind are unreliable — the same problem that makes the removed state/attribute boundary unreliable (§A.5) — so rigidity is dropped as a typing axis consistently, not as a one-off concession.

Consistent with DUL. dul:Concept ⊑ dul:SocialObject ⊑ dul:Object, so an anti-rigid intrinsic sortal falls under dul:Object whether or not it is named separately — a loss of specificity, not an error. It also repairs role: under a single dependence test, filho is unproblematic even though being someone's son is arguably rigid (you do not stop being your father's son when he dies) yet plainly externally dependent.

What is lost. Criança and pessoa carry the same type despite differing in rigidity, and an export to a rigidity-aware ontology (UFO, gUFO) would have to decide Kind vs Phase at export time rather than reading it off the stored type. If that export becomes a requirement, a dedicated phase type can be reinstated for anti-rigid intrinsic sortals without invalidating any existing object assignment — a candidate list can be generated by querying object lemmas whose frames are @entity frames built over life-stage, career-stage or status dimensions.

A.3 No abstract type#

Abstract nouns need no type of their own. DUL's Abstract is not "abstract nouns" — it is value spaces: Region, TimeInterval, SpaceRegion, the space of natural numbers. dul:Object already covers social and cognitive entities (SocialObject, InformationObject, Concept, Description), so lei, ideia, teoria, poema are Objects in DUL already.

The only residue is fact and proposition nouns (o fato de que…, a possibilidade), reified Situations that take event.

Mass nouns need no separate type either: água is anti-unity but a rigid sortal, and unity is not a typing axis (§A.1).

Dot objects (livro = physical • information) are deferred: if co-predication comes to require it, add a boolean refinement on object — physical vs social-informational — not a base type (see also §A.9, item 4).

A.4 Grimshaw's nominal diagnostics#

Grimshaw's complex-event/simple-event/result trichotomy separates nominals by argument structure: whether the noun takes obligatory arguments and aspectual modifiers. Simple event nominals (viagem, corrida, evento) fail those tests while remaining ontologically events, so the tests do not identify type (§4). They describe nominal syntax, not this classification.

A.5 quality and value do not distinguish causal origin#

Three reasons back the claim in §4 that causal origin plays no role in the quality/value classification.

It duplicates the frame. Whether a producing event is on record is a fact about the situation, which the frame layer already carries (the STATIVE schema vs Attributes plus a relation to the producing event). A fact that does not distinguish namespaces, and is already carried by the frame, does not need a type-level home.

It matches the asymmetry already accepted for destruição. Destruição stays event while its frames range over @agentive, @change and @stative; by the same logic quebrado stays value while its frame carries the fact that the value resulted from an event.

There is no reliable test. A distinction between a "resultant" quality (cansaço, force-maintained postura) and a non-resultant one (beleza, inteligência) — the removed state/attribute line — re-lexicalizes the stage-level/individual-level predicate distinction (Carlson 1977; Kratzer 1995), which is gradient, frequently ambiguous within a single predicate, and which the semantics literature declines to treat as an ontological sort.

Consequences: postura, cansaço → quality. Cansado, quebrado, sentado → value. Anti-rigidity and causal origin are read off the frame.

A.6 Theoretical grounding for the procedural types#

The conceptual/procedural split is Relevance Theory's (§2). The nine subtypes are not — no RT publication proposes this inventory.

RT subdivides procedural meaning along two axes not used here: the target of the constraint (explicature, higher-level explicature, or implicature — Blakemore 1987; Wilson & Sperber 1993), and the cognitive subsystem the item triggers (Wilson 2011, 2016). The inventory here is organized by functional category instead, applying Escandell-Vidal & Leonetti's thesis that the semantics of functional categories is procedural. Their own enumeration — discourse markers, sentence-mood marks, quotative and evidential particles, intonation, verb tenses and moods, definite determiners and pronouns, deictic and focusing adverbs, information-structure devices — is a list, not a typology; the nine types coarsen that list into classes an annotator can assign.

Type Nearest published support Status
reference Wilson & Sperber 1993 on pronouns; Escandell-Vidal & Leonetti on definites and deictics RT-aligned
connection Blakemore's core case RT-aligned
focus focusing adverbs in Escandell-Vidal & Leonetti; Iten on concessives and even RT-aligned
interaction Wharton on interjections; Curcó on Spanish particles; quotative and evidential particles in Escandell-Vidal & Leonetti RT-aligned
predication the functional-category thesis derived, not RT
quantification only definite determiners are standardly procedural in RT departure
polarity classical RT gives não/not a logical entry, i.e. conceptual departure
degree degree semantics (Kennedy; and Gärdenfors for the scale) outside RT
modality RT treats only mood and modal particles as procedural RT-aligned after narrowing (§5)

quantification and polarity. Classical RT (Sperber & Wilson 1986, ch. 2) treats logical words as concepts with logical entries — a claim about how an item enters inference. The criterion used here is narrower: does the lemma name something the conceptualization inventory can hold? Não, todo and nenhum have no dul: referent under any construal, and typing them conceptually would require inventing one.

degree is not an RT category. The "region located against a standard" criterion comes from degree semantics, and the dimension/region framing from Gärdenfors (§A.1). It is procedural because it instructs rather than names — the operational test of §2.

modality. Wilson & Sperber (1993) and Ifantidou (1993, 2001) treat sentence adverbials such as certamente and provavelmente as conceptual; Papafragou (2000) treats modal verbs as conceptual. In RT only mood and modal particles are procedural. FNBr's scalar adverbs are value (§5, "modality vs value"), which keeps this row RT-aligned. One residue: RT would also call the modal auxiliaries conceptual, following Papafragou. They are kept procedural because no conceptual type fits them — they name no scale, no event and no relation — and predication would misdescribe them as contentless.

What the inventory still lacks. Sentence mood and type (declarative, interrogative, imperative) has no type of its own — currently spread across reference (quem) and modality, with free-word mood markers landing in interaction. Whether mood deserves its own type depends on how much of it FNBr lexicalizes; mood realized by verb morphology or intonation is not a lemma and is out of scope. Tense and aspect are procedural on most RT accounts but out of scope here for the same reason: bound morphology, not lemmas.

One RT axis considered and declined. Adding target ∈ {explicature, higher-level explicature, implicature} as an orthogonal attribute would let the scheme cite Wilson & Sperber (1993) directly. It is not adopted: the target of a procedural constraint varies with use (the same connective can strengthen or contradict), so the attribute would be a per-token judgement recorded on a per-lemma row, and nothing in §8 consumes it.

A.7 modality vs interaction for evidentials#

Diz que is interaction on the grounds that it points at an information source rather than stating speaker confidence. Some frameworks treat evidentiality as a species of epistemic modality, which would put it with modality; now that modality is narrowed to non-scalar items, the two are closer than before. Open — see §A.9, item 11.

A.8 The cost of splitting adpositions#

Splitting moves work from the lexicon to the parser: the lexicon records a stable fact and the parser resolves a contextual one. That is the right direction, but it is a real cost and should be measured on the adposition tokens specifically, not averaged into overall lemma-resolution accuracy — see §A.9, item 9.

A.9 Open questions#

Recorded so they are not rediscovered as objections.

  1. Validation of role. What inter-annotator agreement does the dependence test achieve on Portuguese data? The two corpus diagnostics in §4 give an external anchor; the number has not yet been measured.
  2. Rigidity is not recorded. See §A.2.
  3. [U] is not used. Löbner's uniqueness feature distinguishes filho (relational) from pai (functional) and predicts determination and bridging behaviour. FNBr currently collapses both to role. Whether [U] earns a column is open.
  4. Dot objects. object does not represent the dual identity of livro (physical • information) that co-predication exposes. Deferred, not solved.
  5. POS × type correlation. How much independent information does the type layer carry over UPOS? One contingency table answers it (§1.5).
  6. The two-sort split has a boundary. Some lemmas plausibly encode both conceptually and procedurally (sempre, só). The current rule assigns one type; whether that is adequate is untested.
  7. Split count. How many lemmas split under §1.4, and how concentrated is the count? A few dozen forms, dominated by adpositions, is the expected shape (§6.2). If splits are instead spread thinly across the content lexicon, the type inventory is drawing a line the lexicon does not observe, and the inventory — not the rule — needs revisiting.
  8. Relatedness between split lemmas is unrecorded. Negação act/operator and pintura activity/canvas are separate rows with no link. Where the relation matters it should be added explicitly at the lexical layer (OntoLex-Lemon vartrans). Whether it matters enough to build is open.
  9. Adposition disambiguation accuracy. Splitting de, a, em, por, com, para and sobre by type (§6.2) moves their disambiguation to the parser, where UPOS cannot help because the candidates share it. Accuracy on adposition tokens should be measured separately, and the government heuristic evaluated on its own. If it underperforms badly, the alternatives are a richer valence-driven disambiguation or a coarser type inventory for adpositions — not a return to per-LU typing.
  10. The modality/value line needs validating, not deciding. The decision is taken (§5): scalar epistemic adverbs are value, talvez and the modal auxiliaries are modality, on whether the lemma names a lexicalized dimension with a gradable counterpart. What is untested is coverage — whether every PB epistemic adverb falls cleanly on one side. Possivelmente and eventualmente are the likely stress cases, and eventualmente may not be epistemic at all in PB. Run the three tests over a closed list of epistemic adverbs before annotation opens.
  11. The modality/interaction boundary for evidentials. See §A.7.
  12. interaction has not been tested at scale. It was added for FNBr's pragmatic frames and its three boundaries — against connection, modality and focus (§5) — are stated but unvalidated. PB interactional particles are frequent and often multifunctional (tá as tag, as agreement token, as imperative of estar), so this type is the most likely source of new §1.4 splits.
  13. The construal-shift diagnostic for substantivized adjectives is new and untested. The "recoverable covert sortal" test (§1.4) separates productive ADJ→N ellipsis (o bonito, um americano) from lexicalized conversion (dependenteₙ, and presumably empregado, acusado and similar deverbal/deadjectival nouns). Coverage across the lexicon is unmeasured. The ~25 "people by X" @entity frames (§7) are themselves a conventionalized middle case — a nationality, religion or vocation adjective defaults to a person referent systematically enough to be frame-licensed, yet no lemma split is made — and it is untested whether other systematic default-referent classes exist that the general test would miss.