Author(s): Quantum Stat Beware the Beautiful Witch I compared various resources for code, research and apps. It turned out that not all NLP research and code are on arXiv. The NLP Index was created to provide a greater range of NLP code and research. Search-as-you’d-type engine that contains over 3 ,000 RSS NLP repositories. It is updated weekly. This index includes the research paper and a ConnectedPapers link to a graph of similar papers. It also contains its GitHub repo. NLP Index This platform allows researchers and hackers access to quick and complete information about NLP. Not just research papers but also the amazing apps created from this research. Because of the interdependencies between subject areas, we have included open search. Sometimes a paper/repo could be about both “knowledge graphs”, and “datasets,” and it can be difficult to separate topics. The user has the choice to openly search the database in any domain/sector simultaneously. For convenience, we also added pre-defined queries that cover dozens of topics within NLP to the sidebar. There are several features to the index, including search as you type and typo tolerance as well as synonym detection. Search for synonyms. For instance, if “dataset” is typed in, the index will search simultaneously for “corpus”, “corpora”, and other text to ensure that every asset has been searched. Typo Tolerance: If you search for “gpt2”, it will include “gpt-2”. Searching in real time will produce results on all characters you enter. It takes only a few milliseconds. Thanks to memory mapping, the Big Bad NLP Database and the NLP Index have been combined! To view the latest compendium NLP datasets you can click the dataset link in the sidebar or search openly for specific tasks/datasets. I’ll eventually remove the BBND URL from the sidebar and redirect it to Index. Thank you for all the help I received since launching the NLP Index. Philip Vollet, thank you for sharing your dataset and hundreds of NLP repos. His posts can be found in the “Uncharted” section. Stay tuned for more features. Keep checking back. Let’s explain ourselves with BERT! Find out why BERT draws an inference with SHAP (SHapley additive exPlanations). This is a game-theoretic method to describe the output of any machinelearning model. This pipeline is used by Transformers. Quick-tips/ml6team Explainable AI Cheat sheet Includes a graphic and YouTube video, as well as links to papers and books on the subject of explainable AI. The StyleCLIP Explainable AI Guide is too much fun! Max Woolf gives a great introduction to StyleCLIP, via Colab notebooks. StyleCLIP allows you to modify headshot photos using text prompts. The quality of the pictures is excellent, and you can also add your photos. For example, take a look at the generation after the text prompt: “Face after using the NLP index” Easily Transform Portraits of People into AI Aberrations Using StyleCLIP | Max Woolf’s Blog Software Updates AdapterHub New version includes BART and GPT-2 models Adapters for Generative and Seq2Seq Models in NLP BERTopic (semi-)supervised topic modeling by leveraging supervised options in UMAP model.fit(docs, y=target_classes) Backends: Added Spacy, Gensim, USE (TFHub) Use a different backend for document embeddings and word embeddings Create your own backends with bertopic.backend.BaseEmbedder Click here for an overview of all new backends Calculate and visualize topics per class Calculate: topics_per_class = topic_model.topics_per_class(docs, topics, classes) Visualize: topic_model.visualize_topics_per_class(topics_per_class) Release Major Release v0.7 MaartenGr/BERTopic Repo Cypher A collection of recently released repos that caught our Gradient-based Adversarial Attacks against Text Transformers A general-purpose framework, GBDA (Gradient-based Distributional Attack), for gradient-based adversarial attacks, and apply it against transformer models on text data. facebookresearch/text-adversarial-attack Connected Papers Easy and Efficient Transformer Pytorch inference plugin for transformers with large model sizes and long sequences. It currently supports GPT-2, and BERT models. NetEaseFuXi/EETConnected papers MDETR: Modulated detection for End-to–End Multi-Modal understanding Code. Links to pre-trained MDETR models (ModulatedDETR) are provided. This allows for fine-tuning tasks that require fine-grained understanding both image and text. ashkamath/mdetrConnected Papers XLM–T — Multilingual Language Model Toolkit For Twitter. Pre-training continues on a large corpus in Twitter in multiple languages using the XLM–Roberta–Base model. Four colab notebooks are included. Cardiffnlp/xlmt Connected papers FRANK: Factuality Assessment Benchmark A typology for factual errors that allows fine-grained analysis and summarization of facts. Artidoro/frankConnected Papers Legal document similarity An assortment of the most advanced methods of representing documents for retrieving US case law that is semantically related. We explored text-based methods (e.g. fastText, Transformers), as well as citation-based (e.g. DeepWalk, Poincare) and hybrid options. malteos/legal-document-similarity Connected Papers Dataset of the Week: Shellcode_IA32 What is it? The Shellcode_IA 32 dataset contains 20 year of shellcodes. It is one of the most comprehensive collections of assembly shellcodes. The instructions in the assembly language of IA -32 are 3 from public security exploits. This dataset is used to automatically generate shell code (code generation job). These assembly programs were used to create shellcode using shell-storm and exploit-db. paper What is the problem? dessertlab/Shellcode_IA32 Every Sunday we do a weekly round-up of NLP news and code drops from researchers around the world. 05.02. 21 originally appeared in on Medium. People are responding and highlighting this story. Published via
Home Innovation 05.02.21
05.02.21
THE FOREFRONT OF TECHNOLOGY
We monitors and writes about new technologies in areas such as technology, innovation, digitization, space, Earth, IT and AI.







