Jokes help humans interact and relate with each other, so why not share a laugh with robots too? Despite the failure of most sci-fi androids to grasp the intricacies of humor, researchers are developing a system to identify jokes and puns. They hope it will help increase machine understanding of common language usage, thus improving human-computer interaction.
via KurzweilAI.net
Showing posts with label nlp. Show all posts
Showing posts with label nlp. Show all posts
Wednesday, August 8, 2007
More Powerset Hype
MIT Technology Review has another article on Palo Alto startup Powerset (previously covered here), which promises to use NLP and semantic technologies to revolutionize search. Nothing new here, just the same promises as before. As always, I’ll remain skeptical until I see the product in action.
By the way, I’m still waiting for my beta invitation.
via KurzweilAI.net
By the way, I’m still waiting for my beta invitation.
via KurzweilAI.net
Thursday, July 26, 2007
Baby Talk
Stanford researchers are working on an interesting alternative to building natural language rules by hand... having the software learn the language on it's own like a human child. The idea is for the system to analyze and sort through speech sounds until it understands language structure. While I agree that it will be much easier to build a system that can learn and acquire language on its own, it will need to be "seeded" with some general rules of grammar, much like the innate rules that some believe human babies are born with.
via AI-Depot.com
via AI-Depot.com
Google’s Future
MIT Technology Review recently posted an interview with Peter Norvig, director of research at Google, regarding the future of search. An AI expert, Norvig sees machine translation and speech recognition as the next big things to improve Google's search and advertising. He also identifies understanding the contents of documents as one of the two biggest problems in search... leading to much NLP work ongoing at Google.
via AAAI.org News
via AAAI.org News
Close, But Not Quite
Slashdot reported a researcher has created a text compression program capable of coming within 1% of the AI equivalent of human performance as determined by Claude Shannon.
What does this mean? Are we 99% of the way to “achieving” AI? No, it simply means we have AI tools that are 99% as capable as humans when it comes to text compression. We already have computers that are better at chess than humans, so this is simply another domain where algorithms are successfully competing with neurons.
What does this mean? Are we 99% of the way to “achieving” AI? No, it simply means we have AI tools that are 99% as capable as humans when it comes to text compression. We already have computers that are better at chess than humans, so this is simply another domain where algorithms are successfully competing with neurons.
Friday, June 29, 2007
How would NLP parse buzzwords?
More about Powerset (covered here before), which Techdirt seems to think is little more than buzzwords and patent threats. The start-up claims to be developing natural language search technology, and recently held an event in San Francisco to unveil itself to the world.
Among its many lofty goals, Powerset wants to become the ultimate web system, by creating what ZDNet calls "the natural language search mashup platform." For now, I've got to be as skeptical as Techdirt and think that these folks just combined as many hot buzzwords as they could come up with and slapped a couple of questionable patents on them. This kind of talk is a great way to generate venture capital funding, but likely won't do much to advance NLP. Hopefully it will turn out that Powerset has something great in store for all of us, and that this is all just some marketing and PR run amok, but until then we'll just have to wait & see.
Among its many lofty goals, Powerset wants to become the ultimate web system, by creating what ZDNet calls "the natural language search mashup platform." For now, I've got to be as skeptical as Techdirt and think that these folks just combined as many hot buzzwords as they could come up with and slapped a couple of questionable patents on them. This kind of talk is a great way to generate venture capital funding, but likely won't do much to advance NLP. Hopefully it will turn out that Powerset has something great in store for all of us, and that this is all just some marketing and PR run amok, but until then we'll just have to wait & see.
Tuesday, June 26, 2007
NLP to Earn Big Bucks
The NY Times covers news mining algorithms designed to digest the volumes of information available on the Internet in news articles, journals, studies, and legal filings. Financial institutions are using these programs to generate massive amounts of stock trades, easily replacing large staffs of analysts. Much like reactive day-traders who launch waves of trades based on buzzwords found in headlines, these systems look for key words & phrases known to be trade triggers to predict movement of individual stocks and sectors.
Monday, June 25, 2007
Overrated Semantic Search
Several IT news outlets have been fawning over Xerox's new semantic-based search engine, which I've covered before. The general idea behind the technology is to analyze linguistic structures in order to improve search results.
Xerox plans to use this technology in legal software to enable "e-discovery" by sifting through massive amounts of documents, searching for information relevant to a case. Perhaps this will lead to the second instance of software being sued for practicing law without a license.
This is all well and good, and a natural progression for the science of search. Not all that dramatic an improvement as the articles would lead us to believe, buy hey, you gotta sell papers, right? However, it aggravates me when the media makes a factual error while covering a topic I’m familiar with...
Except a Google search for “lincoln’s first vice president” provides the correct result, as does running “Who was Lincoln’s first vice president?” through my quite unsophisticated Answer Machine. While I can’t fault the reporter for overlooking my humble research, shouldn’t they be capable of running a simple Google query? Wouldn’t this fall under the category of “thorough fact checking?” Shouldn’t they run their “facts” through a subject matter expert before publishing them? And I mean an actual expert, not a PR staffer from the company at hand. Unfortunately, more and more tech articles in the media have regressed to little more than paraphrasing press releases.
Furthermore, if I notice obvious mistakes regarding topics I know a little something about, what other incorrect information am I obliviously consuming? And the media wonders why we don’t trust them anymore...
Xerox plans to use this technology in legal software to enable "e-discovery" by sifting through massive amounts of documents, searching for information relevant to a case. Perhaps this will lead to the second instance of software being sued for practicing law without a license.
This is all well and good, and a natural progression for the science of search. Not all that dramatic an improvement as the articles would lead us to believe, buy hey, you gotta sell papers, right? However, it aggravates me when the media makes a factual error while covering a topic I’m familiar with...
For example, common searches using keywords "Lincoln" and "vice president" likely won't reveal President Abraham Lincoln's first vice president. A semantic search should yield the answer: Hannibal Hamlin.
Except a Google search for “lincoln’s first vice president” provides the correct result, as does running “Who was Lincoln’s first vice president?” through my quite unsophisticated Answer Machine. While I can’t fault the reporter for overlooking my humble research, shouldn’t they be capable of running a simple Google query? Wouldn’t this fall under the category of “thorough fact checking?” Shouldn’t they run their “facts” through a subject matter expert before publishing them? And I mean an actual expert, not a PR staffer from the company at hand. Unfortunately, more and more tech articles in the media have regressed to little more than paraphrasing press releases.
Furthermore, if I notice obvious mistakes regarding topics I know a little something about, what other incorrect information am I obliviously consuming? And the media wonders why we don’t trust them anymore...
Tuesday, June 5, 2007
Using NLP to Organize Unstructured Data
Another facet of the information overload problem is trying to get a handle of the volumes of unstructured data created by organizations on a daily basis, and package them into a searchable, manageable package. Some establishment struggle with file plan compliance and enforcement, while others provide tools to index and search documents based on keywords. IBM, on the other hand, is applying NLP techniques to try and solve the problem. OmniFind tackles content classification by scanning varied types of unstructured data, automatically learning and categorizing information into newly-created as well as existing taxonomies. By understanding linguistics, semantics, and context, OmniFind is able to determine connections and make inferences beyond the reach of even the greatest keyword-based search algorithms. Another example of NLP making information easier to find, access, and use.
Stovepipe NLP Research
The National Science Foundation is sponsoring research into NLP designed to help government clerks get a handle of the information overload coming from the glut of public comments pouring into www.regulations.gov. The site allows officials to solicit and consider public comments while creating rule and regulations concerning things like organic food labeling and media ownership consolidation. It seems to be a success, as far more comments are submitted than can be effectively sorted through by hand. While it seems reasonable to apply NLP techniques to this problem, should the research money be directed at something like the more general problem of information overload than such a narrow application as this?
Thursday, May 31, 2007
Semantic Search is Coming
I came across an article about semantic search that does a decent job of explaining the differences between statistics-based and semantics-based approaches to information retrieval. It also describes some of the difficulties and shortcomings of the Semantic Web, along with a few of the related natural language applications.
via Slashdot
via Slashdot
Wednesday, April 25, 2007
NLP Calculator for Mac
The Unofficial Apple Weblog points us in the direction of a calculator that can solve a limited scope of natural language word problems. Written as a front-end for the GNU bc calculator, Soulver is something fun to play with but isn't meant to satisfy all your calculating needs.
I was recently mulling around the idea of creating a natural language Unix shell. It wouldn't be too difficult to map out a significant number of input patterns to actual Unix/Linux commands, but the necessarily narrow scope of the system would have been a major usability drawback, as it must be for Soulver.
I was recently mulling around the idea of creating a natural language Unix shell. It wouldn't be too difficult to map out a significant number of input patterns to actual Unix/Linux commands, but the necessarily narrow scope of the system would have been a major usability drawback, as it must be for Soulver.
Monday, April 2, 2007
Google Speaks 12 Languages
I just read an interesting article about the great success Google has been having using statistical machine translation for automatic translation of foreign language documents, as opposed to the rule & grammar based approaches used before.
I, for one, am very much persuaded by the idea that human language is so complex (we don't even fully understand it, just ask a linguist!) that it can never be fully hard-coded into a machine, but rather, a machine must "learn" it on it's own for the most part (we'll guide it and help it where needed). I like the approach that Google is taking, but without a conceptual framework or knowledge representation model, the system really isn't "understanding" or "comprehening" anything--it's just doing a "dumb" translation using statistical references. Still quite an accomplishment, but entirely different from my ultimate objective.
I, for one, am very much persuaded by the idea that human language is so complex (we don't even fully understand it, just ask a linguist!) that it can never be fully hard-coded into a machine, but rather, a machine must "learn" it on it's own for the most part (we'll guide it and help it where needed). I like the approach that Google is taking, but without a conceptual framework or knowledge representation model, the system really isn't "understanding" or "comprehening" anything--it's just doing a "dumb" translation using statistical references. Still quite an accomplishment, but entirely different from my ultimate objective.
Tuesday, March 20, 2007
Machine Analysis of Scientific Papers
There's a lot of exciting work going on in NLP right now, and it's hard to keep up...and even harder to maintain a blog about all of it! Larry pointed me to an article from last year detailing an automated tool for analyzing and comparing experimental reports.
This project sounds like some sort of XML markup scheme for outlining scientific papers, similar to the ontologies powering the semantic web initiative. It probably involves too much overhead to be widely adopted and therefore be useful, as it likely requires the author to spend an awful lot of additional time constructing papers such that the EXPO system could parse it. A much more elegant method would be for the system to perform an automatic analysis and markup of the text, however that would require NLP technology beyond what's currently available.
As you might imagine, a similar hurdle exists for the adoption of the semantic web in general. However, the analysis & synthesis of peer-reviewed journals presents us with yet another "killer app" for NLP. There is simply far too much information covering any given topic being generated for a human being to digest, even experts in a particular field...let alone a renaissance man or polymath to synthesize from diverse fields. Only a machine with advanced NLP capabilities would be able to make sense of it all and create new knowledge from what's already available.
This project sounds like some sort of XML markup scheme for outlining scientific papers, similar to the ontologies powering the semantic web initiative. It probably involves too much overhead to be widely adopted and therefore be useful, as it likely requires the author to spend an awful lot of additional time constructing papers such that the EXPO system could parse it. A much more elegant method would be for the system to perform an automatic analysis and markup of the text, however that would require NLP technology beyond what's currently available.
As you might imagine, a similar hurdle exists for the adoption of the semantic web in general. However, the analysis & synthesis of peer-reviewed journals presents us with yet another "killer app" for NLP. There is simply far too much information covering any given topic being generated for a human being to digest, even experts in a particular field...let alone a renaissance man or polymath to synthesize from diverse fields. Only a machine with advanced NLP capabilities would be able to make sense of it all and create new knowledge from what's already available.
Friday, March 2, 2007
Commercially Available Automatic Summarization Software
In his weekly article, Robert X. Cringley profiles a company that has created software which is seemingly able to create relevant summaries of arbitrary size from bodies of text covering all possible subject domains. This is precisely the sort of thing I wanted to accomplish with my AutoSummary project. Automatic summarization is yet another application of NLP as a solution to the problem of too much information for humans to deal with. Learn more about iReader at Syntactica.
Saturday, February 24, 2007
"Text Enrichment" to Improve Written English
An interesting application of NLP compares user-generated input against a vast database of known proper English to generate suggested improvements to help readability and fluency.
By using a corpus filled with millions of real-world modern English texts, the software is capable of recommending thousands of grammatical corrections, as well as relevant adjectives, adverbs and synonyms. The Israel-based company WhiteSmoke hopes to help improve emails and documents by leaping well ahead of the limited grammar checking functions found in programs like Microsoft Word.
By using a corpus filled with millions of real-world modern English texts, the software is capable of recommending thousands of grammatical corrections, as well as relevant adjectives, adverbs and synonyms. The Israel-based company WhiteSmoke hopes to help improve emails and documents by leaping well ahead of the limited grammar checking functions found in programs like Microsoft Word.
Saturday, February 17, 2007
PARC to build NLP Search Engine
The Palo Alto Research Center (of Xerox fame) has licensed its sophisticated natural language processing technology to a start-up hoping to develop an NLP-powered search engine.
The start-up, called Powerset, intends to create a system where users search for data by entering plain-language queries, rather than using keywords. Similar efforts have been launched by MIT and others to solve the problems of automated response generation. With decades of research and significant resources behind them, Powerset hopes to foster the third generation of search engines, following in the footsteps of AltaVista and Google.
The start-up, called Powerset, intends to create a system where users search for data by entering plain-language queries, rather than using keywords. Similar efforts have been launched by MIT and others to solve the problems of automated response generation. With decades of research and significant resources behind them, Powerset hopes to foster the third generation of search engines, following in the footsteps of AltaVista and Google.
Monday, January 29, 2007
Inside the Walls of DARPA
Computerworld offers A Peek Inside DARPA, the Defense Advanced Research Projects Agency responsible for inventing the Internet and other advanced technologies.
Projects of note include Global Autonomous Language Exploitation (GALE), which is designed to "transcribe, translate & distill" data collected from English and foreign language sources into actionable information for human decision makers. Another, the Personalized Assistant that Learns (PAL), is hoped will one day "automatically watch a conversation between two people and, using natural-language processing, figure out what are the tasks they agreed upon."
GALE in particular seems to be a promising endeavor. Researchers at Penn's Linguistic Data Consortium (LDC) are working to provide resources and tools for the project. Possibly the most interesting and impactful component is the Distillation engine, which seems to be the segment focused on "understanding" the information at hand. LDC has posted a thorough background & analysis of this function on their website.
Projects of note include Global Autonomous Language Exploitation (GALE), which is designed to "transcribe, translate & distill" data collected from English and foreign language sources into actionable information for human decision makers. Another, the Personalized Assistant that Learns (PAL), is hoped will one day "automatically watch a conversation between two people and, using natural-language processing, figure out what are the tasks they agreed upon."
GALE in particular seems to be a promising endeavor. Researchers at Penn's Linguistic Data Consortium (LDC) are working to provide resources and tools for the project. Possibly the most interesting and impactful component is the Distillation engine, which seems to be the segment focused on "understanding" the information at hand. LDC has posted a thorough background & analysis of this function on their website.
Thursday, October 26, 2006
Max Headroom Lives!
Coming soon to your TV: newscasts written, produced and anchored entirely by autonomous software.
Think Ananova without the writers. The system, called News at Seven, generates a script from RSS news feeds customized to a viewer’s interests which is then read by an avatar created using the Half-Life game engine. The system can also augment the broadcast with clips from YouTube or other video sites based on keywords in the news stories.
News at Seven was created by a team of researchers at the Northwestern University Intelligent Information Lab led by Kristian Hammond along with grad students Nathan Nichols and Sara Owsley. Several videos are already available for viewing, including a report on the alleged North Korean nuclear test.
A novelty for now, perhaps, but possibly the next step in the long journey towards building systems we interact with in a much more natural way. Imagine an avatar you could ask questions of or hold a conversation with. Using methods such as those described above, this avatar could instantly become an expert in almost any conceivable subject. Or, instead of re-formatting RSS feeds for a news script, the system could follow some sort of knowledge representation framework to build an information bank for later use. If we replace the RSS feeds with some other form of input, say sensory input, it could create “memories” of “personal experiences.” The question at the heart of the matter is: how do we define understanding? Does the News at Seven avatar understand the stories it presents to its audience? Most would agree that it does not. But what if the avatar was able to keep a record of the information from these stories in an accessible memory bank, and could discuss the matter with other people (or avatars) to formulate decisions, actions or even opinions based on that knowledge? Would that qualify as understanding? Or intentional behavior? Or even conscious thought? What do we humans, as conscious beings, do beyond this that leads us to define it all as consciousness and understanding?
via Slashdot
Think Ananova without the writers. The system, called News at Seven, generates a script from RSS news feeds customized to a viewer’s interests which is then read by an avatar created using the Half-Life game engine. The system can also augment the broadcast with clips from YouTube or other video sites based on keywords in the news stories.
News at Seven was created by a team of researchers at the Northwestern University Intelligent Information Lab led by Kristian Hammond along with grad students Nathan Nichols and Sara Owsley. Several videos are already available for viewing, including a report on the alleged North Korean nuclear test.
A novelty for now, perhaps, but possibly the next step in the long journey towards building systems we interact with in a much more natural way. Imagine an avatar you could ask questions of or hold a conversation with. Using methods such as those described above, this avatar could instantly become an expert in almost any conceivable subject. Or, instead of re-formatting RSS feeds for a news script, the system could follow some sort of knowledge representation framework to build an information bank for later use. If we replace the RSS feeds with some other form of input, say sensory input, it could create “memories” of “personal experiences.” The question at the heart of the matter is: how do we define understanding? Does the News at Seven avatar understand the stories it presents to its audience? Most would agree that it does not. But what if the avatar was able to keep a record of the information from these stories in an accessible memory bank, and could discuss the matter with other people (or avatars) to formulate decisions, actions or even opinions based on that knowledge? Would that qualify as understanding? Or intentional behavior? Or even conscious thought? What do we humans, as conscious beings, do beyond this that leads us to define it all as consciousness and understanding?
via Slashdot
Monday, February 13, 2006
Frustrating Recursion
I thought I made a mistake once, but I was wrong. For some reason, this happens to me all the time when programming.
I was working on AutoSummary this weekend, adding a contextual framework using hypernym (superordinate) information for individual senses of a given word. In order to do this I needed to create a b-tree data structure. I thought I had set everything up properly, except the tree wouldn't populate past the second level. Weird crashes, etc. I tried everything, checked all of the functions and methods, testing everything I could think of. Nothing.
Tonight, I started checking everything over again, trying some different approaches. I did some digging and determined something was generating a null pointer exception. I checked everything all over again, and again... nothing. After a period of insufferable aggrivation, I discovered by trial and error that the exception was caused by the fact that I had forgotten to initialize the data container (an ArrayList). I was so worried about getting all the "hard stuff" figured out that I had overlooked a beginner's error.
The moral of the story? The simplest answers are often the hardest to find.
I was working on AutoSummary this weekend, adding a contextual framework using hypernym (superordinate) information for individual senses of a given word. In order to do this I needed to create a b-tree data structure. I thought I had set everything up properly, except the tree wouldn't populate past the second level. Weird crashes, etc. I tried everything, checked all of the functions and methods, testing everything I could think of. Nothing.
Tonight, I started checking everything over again, trying some different approaches. I did some digging and determined something was generating a null pointer exception. I checked everything all over again, and again... nothing. After a period of insufferable aggrivation, I discovered by trial and error that the exception was caused by the fact that I had forgotten to initialize the data container (an ArrayList). I was so worried about getting all the "hard stuff" figured out that I had overlooked a beginner's error.
The moral of the story? The simplest answers are often the hardest to find.
Subscribe to:
Posts (Atom)