Semester project 3 for Linguistics
- analyze some text
- learn techniques for analyzing text
- learn some software engineering skills
- learn
git,markdown, useLaTeX
-
how many tasks? 4
-
how many teams? 4
-
Introduce yourselves
- name
- languages spoken (natural/programming)
- interests (relevant to the course)
- one fun fact (not relevant to the course)
-
Body-focused repetitive behaviour (BFRB) Sara
- look at Doctor-patient interaction
- look for examples of BFRB in literature/media
- TASK:
- READ A Qualitative Study Exploring International Experiences of Seeking Treatment for Adults With Trichotillomania: A Story of Frustration and Unmet Need
- READ Self-control and Body-focused Repetitive Behaviors
- look at Books about Trichotillomania and Dermatillomania could we analyze these? Are there similar books in Czech?
-
How women are represented online. Anna
- For example in news articles and user comments and maybe comparing different time periods to see how the language changed over time.
- I would also like to compare how men and women talk about each other online.
- TASK: find some literature, refine the problem
-
Distribution of lexicalized metaphor Erika
- combine sense-tagged text with ChainNet
- analyze distribution
- fill in what is missing
- TASK: Read these papers:
-
Improve Czech wordnet Zuzana
- fill in Czech specific patterns
- adverbs [of languages]
- diminutives
- aspect variants
- gender variants (role nouns)
- generate definitions for derived entries
- get examples from corpora
- derivational links with affixes?
- READ Papers:
- Czech Wordnet 1.9 (note we have a private copy of Czech Wordnet 2.1 from Adam Rambousek).
- Derivational Relations in Czech WordNet
- Overview and Future of Czech Wordnet (note the open release did not happen)
- Lexico-Semantic Annotation of the Prague Dependency Treebank
- https://deb.fi.muni.cz:8005/debdict/
- https://cygnet.maudslay.eu/
- https://pypi.org/project/wn/
- fill in Czech specific patterns
-
LLM text vs Human text
- what is different (test for Czech?)
- I have data for English
- TASK: Read these papers:
- get something useful for each task
- combine to make a best-of-breed
- write, submit and publish a paper
- release at least one automatically tagged, aligned corpus
- use github to coordinate
-
Erika and Anna Monday
-
Sara and Zuzana Wednesday (3pm)
-
15-30 minutes progress
-
longer discussion of issues as necessary
- make github account
- send me accountname
- I will add to github
- Background reading, refine problems