core/oga - oga

Commit Graph

Author	SHA1	Message	Date
Yorick Peterse	8fbc582547	XPath evaluation for name/namespace wildcards.	2014-07-09 22:09:20 +02:00
Yorick Peterse	ed45058983	Basic support for evaluating XPath wildcards.	2014-07-09 20:06:31 +02:00
Yorick Peterse	74130df40d	XPath evaluation support for absolute paths.	2014-07-09 09:28:00 +02:00
Yorick Peterse	a398af19fe	Use a stack based XPath evaluator. This evaluator is capable of processing very, very basic XPath expressions. It does not yet know anything about axes, functions, operators, etc.	2014-07-08 23:57:21 +02:00
Yorick Peterse	9c661e1e60	Added XML::NodeSet#+ and XML::NodeSet#to_a	2014-07-08 23:25:09 +02:00
Yorick Peterse	266c66569e	Removed internal state of XPath::Evaluator. Come to think of it it might actually be easier to implement the evaluator as an actual VM. That is, instead of directly running on the AST it runs on some flavour of bytecode. Alternatively it runs directly on the AST but behaves more like a (stack based) VM. This would most likely be easier than passing a cursor to every node processing method.	2014-07-08 19:19:59 +02:00
Yorick Peterse	808e1e8c47	Initial, half-assed attempt at an XPath evaluator.	2014-07-07 19:41:09 +02:00
Yorick Peterse	8b381ac970	Added Node#remove. This method can be used to remove individual nodes without first having to retrieve the NodeSet they are stored in.	2014-07-04 10:26:41 +02:00
Yorick Peterse	e334e50ca6	Added Node#previous_element and Node#next_element. These methods can be used similar to #previous and #next expect that they only return Element instances opposed to all Node instances.	2014-07-04 10:18:18 +02:00
Yorick Peterse	94965961ce	Added XML::Element#inner_text. This method can be used to retrieve the text of the given node only. In other words, unlike Element#text it does not contain the text of any child nodes.	2014-07-04 09:50:47 +02:00
Yorick Peterse	a15c19c4af	Retrieving root nodes using Node#root_node. This method uses a loop to traverse upwards the DOM tree in order to find the root document/element. While this might have an impact on performance I don't expect Oga itself to call this method very often. The benefit is that Node instances don't require users to manually pass the top level document as an argument.	2014-07-03 00:26:15 +02:00
Yorick Peterse	e69fbc3ea7	Remove ownership when using NodeSet#delete.	2014-07-01 10:15:39 +02:00
Yorick Peterse	4e2b7f529d	Cleared up docs a bit of XML::Node.	2014-07-01 10:03:04 +02:00
Yorick Peterse	30845e3d65	Reliably remove nodes from owned sets. The combination of iterating over an array and removing values from it results in not all elements being removed. For example: numbers = [10, 20, 30] numbers.each do \|num\| numbers.delete(num) end numbers # => [20] As a result of this the NodeSet#remove method uses two iterations: 1. One iteration to retrieve all NodeSet instances to remove nodes from. 2. One iteration to actually remove the nodes. For the second iteration we iterate over the node sets and then over the nodes. This ensures that we always remove all nodes instead of leaving some behind.	2014-07-01 09:57:43 +02:00
Yorick Peterse	5073056831	Added require() for StringIO. This is needed since the lexer checks if the input is a StringIO instance.	2014-07-01 09:37:52 +02:00
Yorick Peterse	8314d24435	First pass at using the NodeSet class. The documentation still leaves a lot to be desired and so does the API. There also appears to be a problem where NodeSet#remove doesn't properly remove all nodes from a set. Outside of that we're making slow progress towards a proper DOM API.	2014-06-30 23:03:48 +02:00
Yorick Peterse	e71fe3d6fa	Indexing of NodeSet instances. Similar to arrays the NodeSet class now allows one to retrieve nodes for a given index.	2014-06-29 23:59:27 +02:00
Yorick Peterse	253575dc37	Basic docs for the NodeSet class.	2014-06-29 21:34:18 +02:00
Yorick Peterse	d2e74d8a0b	Specs for the NodeSet class.	2014-06-26 19:52:03 +02:00
Yorick Peterse	eb9d4fbccc	Changed NodeSet to behave more like an Array.	2014-06-26 09:37:54 +02:00
Yorick Peterse	a98f50b63b	NodeSet#push should not take node ownership.	2014-06-26 09:37:30 +02:00
Yorick Peterse	15fa7a2068	Remove explicit index tracking of NodeSet. Instead the various nodes can use NodeSet#index (aka Array#index) instead. This has a slight performance overhead on very large (millions) of nodes but should be fine in most other cases.	2014-06-25 09:41:58 +02:00
Yorick Peterse	884dbd9563	Rough sketch for a NodeSet class.	2014-06-24 19:06:45 +02:00
Yorick Peterse	b8cc6b5031	XPath::Parser#parse returns a XPath::Node instance	2014-06-23 20:23:08 +02:00
Yorick Peterse	03be9f241c	Docs on Racc operator precedence.	2014-06-23 10:36:42 +02:00
Yorick Peterse	96a7a40fdc	Disable code coverage for JRuby specific code.	2014-06-23 09:42:14 +02:00
Yorick Peterse	d2f15e37d0	Corrected XPath operator precedence. The previous commit didn't fully change the operator precedence according to the XPath 1.0 specification. Also thanks to @whitequark for clearing up a few things about Racc's operator precedence system.	2014-06-23 00:30:42 +02:00
Yorick Peterse	a440d3f003	Fixed XPath operator precedence. Apparently using multiple `left` rules with T_AND and T_OR being separate solves this problem. Riiiiight....	2014-06-23 00:15:43 +02:00
Yorick Peterse	6a2f4fa82d	Parsing support for more XPath operators. This still messes up some tests due to botched token precedence (by the looks of it).	2014-06-22 21:27:53 +02:00
Yorick Peterse	514c342cab	Lex wildcards as T_IDENT instead of T_STAR.	2014-06-20 20:37:34 +02:00
Yorick Peterse	45db337c76	Use Array#unshift for multiple xpath call args.	2014-06-17 20:12:25 +02:00
Yorick Peterse	b3ffc28cc7	Removed shift/reduce conflict in the xpath parser.	2014-06-17 20:09:44 +02:00
Yorick Peterse	497f57ccd2	Basic parser setup for XPath function calls.	2014-06-17 19:57:17 +02:00
Yorick Peterse	894de7f909	Lex all XPath expressions in a single machine. This allows literal values such as strings and numbers to be used as function arguments.	2014-06-17 19:56:57 +02:00
Yorick Peterse	bb7af98257	Updated used ASTs for all XPath parser specs.	2014-06-17 18:51:33 +02:00
Yorick Peterse	2298ef618b	Reworked handling of relative vs absolute XPaths.	2014-06-16 20:19:39 +02:00
Yorick Peterse	eba2d9954d	Support for parsing basic XPath expressions.	2014-06-12 00:20:46 +02:00
Yorick Peterse	70f3b7fa92	Lex XPath operators using individual tokens. Instead of lexing every operator as T_OP they now use individual tokens such as T_EQ and T_LT.	2014-06-09 23:35:54 +02:00
Yorick Peterse	7244e28eec	Corrected docs of the xpath lexer.	2014-06-04 19:33:46 +02:00
Yorick Peterse	1d2f9e6db6	Added T_STAR as an XPath parser token.	2014-06-02 09:27:25 +02:00
Yorick Peterse	e11b9ed32c	Basic XPath parser setup.	2014-06-01 23:02:28 +02:00
Yorick Peterse	54de2df0c7	Support for lexing XPath wildcard expressions. To support this we need to require whitespace around the "*" operator. This is not ideal but it will do for now.	2014-06-01 23:01:24 +02:00
Yorick Peterse	8dd8d7a519	Basic working XPath lexer. This doesn't lex everything of the XPath specification just yet and needs more tests.	2014-06-01 19:24:35 +02:00
Yorick Peterse	a50b76a2d8	Cleaned up XPath lexer boilerplate a bit.	2014-05-29 19:25:49 +02:00
Yorick Peterse	e0b07332d9	Boilerplate for the XPath lexer.	2014-05-29 19:25:49 +02:00
Yorick Peterse	be3f8fb494	Removed the on_newline XML lexer callback.	2014-05-29 14:21:48 +02:00
Yorick Peterse	ead5c71fee	Cleaned up the XML parser grammar. This resolves all shift/reduce and reduce/reduce conflicts that were previously present.	2014-05-29 01:37:19 +02:00
Yorick Peterse	49780e2b04	Fix for useless XML parser rules. Something tells me that using : and \| in your syntax might not be the best decision.	2014-05-28 21:36:06 +02:00
Yorick Peterse	28edc7726f	Rewind IO input upon resetting the lexer.	2014-05-26 00:33:20 +02:00
Yorick Peterse	629dcd3fe6	Support for IO inputs in the lexer. Using IO/StringIO objects one can parse large XML files without first having to read the entire file into memory. This can potentially save a lot of memory at the cost of a slightly slower runtime. For IO like instances the lexer will consume the input line by line. If a String is given it's consumed as a whole instead. A small side effect of reading the input line by line is that text such as "foo\nbar" will be lexed as two tokens instead of one. Fixes #19.	2014-05-26 00:30:39 +02:00
Yorick Peterse	6b9d65923a	Use a method for getting input in the XML lexer. Instead of directly accessing the `data` instance variable the C/Java code now uses the method `read_data`. This is part of one of the various steps required to allow Oga to read data from IO like instances. It also means I can freely change the name of the instance variable without also having to change the C/Java code.	2014-05-21 00:27:23 +02:00
Yorick Peterse	cd0f3380c4	Merge multiple CDATA tokens into a single token. The tokens T_CDATA_START, T_TEXT and T_CDATA_END have been merged together into T_CDATA.	2014-05-19 09:36:19 +02:00
Yorick Peterse	a4fb5c1299	Merge multiple comment tokens into a single one. The tokens T_COMMENT_START, T_TEXT and T_COMMENT_END have been merged into a single token: T_COMMENT. This simplifies both the lexer and the parser.	2014-05-19 09:30:30 +02:00
Yorick Peterse	c891dd88cb	Removed useless code from the XML parser.	2014-05-18 23:30:26 +02:00
Yorick Peterse	81a81f0ab0	Don't create Arrays when not needed.	2014-05-16 17:05:42 +02:00
Yorick Peterse	fd2f727183	Only set explicit ivars in the lexer.	2014-05-15 19:48:18 +02:00
Yorick Peterse	44bf1dd1ca	Split up handling of element names/namespaces. This is now split up on Ragel level, simplifying the corresponding Ruby code.	2014-05-15 10:22:05 +02:00
Yorick Peterse	723a273e4f	Enforce symbols for element attributes. This comes with a little bit of memory overhead but this should be minor in most cases.	2014-05-15 01:04:26 +02:00
Yorick Peterse	f4b9bbd4ac	Removed lazy way of setting instance variables. This process is quite a bit slower compared to setting instance variables directly.	2014-05-15 00:43:13 +02:00
Yorick Peterse	19f04f98f7	Support for lexing/parsing inline doctypes.	2014-05-10 00:28:11 +02:00
Yorick Peterse	fe74d60138	Manually bootstrap JRuby after all. After discussing this with @headius I've decided to do this the manual way anyway. Apparently the basic load service stuff is deprecated and not very reliable.	2014-05-07 22:32:34 +02:00
Yorick Peterse	b8efed5177	Renamed on_start_doctype to on_doctype_start.	2014-05-06 23:18:44 +02:00
Yorick Peterse	2053018d07	Slap JRuby so that it can load the .jar file.	2014-05-06 20:45:26 +02:00
Yorick Peterse	6e685378e0	Setup Ragel for JRuby and load things the hard way	2014-05-06 19:06:04 +02:00
Yorick Peterse	ee756037e7	Removed unused YARD tag.	2014-05-05 09:45:10 +02:00
Yorick Peterse	aeab885a7f	Docs for the Ruby part of the XML lexer.	2014-05-05 09:44:35 +02:00
Yorick Peterse	2689d3f65a	Initial setup using a C extension. While I've tried to keep Oga pure Ruby for as long as possible the performance of Ragel's Ruby output was not worth the trouble. For example, lexing 10MB of XML would take 5 to 6 seconds at least. Nokogiri on the other hand can parse that same XML into a DOM document in about 300 miliseconds. Such a big performance difference is not acceptable. To work around this the XML/HTML lexer will be implemented in C for MRI/Rubinius and Java for JRuby. For now there's only a C extension as I haven't read up yet on the JRuby API. The end goal is to provide some sort of Ragel "template" that can be used to generate the corresponding C/Java extension code. This would remove the need of duplicating the grammar and associated code. The native extension setup is a hybrid between native and Ruby. The raw Ragel stuff happens in C/Java while the actual logic of actions happens in Ruby. This adds a small amount of overhead but makes it much easier to maintain the lexer. Even with this extra overhead the performance is much better than pure Ruby. The 10MB of XML mentioned above is lexed in about 600 miliseconds. In other words, it's 10 times faster.	2014-05-05 00:31:28 +02:00
Yorick Peterse	baaa24a760	Indentation fix in the lexer.	2014-05-04 18:06:43 +02:00
Yorick Peterse	f18e8893de	Removed the buffering crap from the lexer.	2014-05-04 17:39:08 +02:00
Yorick Peterse	9dfdefee47	Removed XML::Lexer#buffering? Instead of wrapping a predicate method around the ivar we'll just access it directly. This reduces average lexing times in the big XML benchmark from 7,5 to ~7 seconds.	2014-05-01 22:59:56 +02:00
Yorick Peterse	f607cf50dc	Use local variables for Ragel. Instead of using instance variables for ts, te, etc we'll use local variables. Grand wizard overloard @whitequark suggested that this would be quite a bit faster, which turns out to be true. For example, the big XML lexer benchmark would, prior to this commit, complete in about 9 - 9,3 seconds. With this commit that hovers around 8,5 seconds.	2014-05-01 13:00:29 +02:00
Yorick Peterse	83f6d5437e	Contextual pull parsing. This adds the ability to more easily act upon specific node types and nestings when using the pull parsing API. A basic example of this API looks like the following (only including relevant code): parser.parse do \|node\| parser.on(:element, %w{people person}) do people << {:name => nil, :age => nil} end parser.on(:text, %w{people person name}) do people.last[:name] = node.text end parser.on(:text, %w{people person age}) do people.last[:age] = node.text.to_i end end This fixes #6.	2014-04-29 23:05:49 +02:00
Yorick Peterse	1a413998a3	Track the current node in the pull parser. The current node is tracked in the instance method `node`.	2014-04-29 21:21:05 +02:00
Yorick Peterse	45b0cdf811	Track element name nesting in the pull parser. Tracking the names of nested elements makes it a lot easier to do contextual pull parsing. Without this it's impossible to know what context the parser is in at a given moment. For memory reasons the parser currently only tracks the element names. In the future it might perhaps also track extra information to make parsing easier.	2014-04-28 23:40:36 +02:00
Yorick Peterse	030a0068bd	Basic pull parsing setup. This parser extends the regular DOM parser but instead delegates certain nodes to a block instead of building a DOM tree. The API is a bit raw in its current form but I'll extend it and make it a bit more user friendly in the following commits. In particular I want to make it easier to figure out if a certain node is nested inside another node.	2014-04-28 17:22:17 +02:00
Yorick Peterse	fd5bbbc9a2	Move element recursion handling into a method. This makes it easier to disable later on in the streaming parser.	2014-04-28 10:25:05 +02:00
Yorick Peterse	785ec26fe7	Create Element instances before recursing.	2014-04-28 10:21:34 +02:00
Yorick Peterse	9939cf49eb	Move parser callback code into dedicated methods.	2014-04-28 10:18:55 +02:00
Yorick Peterse	5d05aed6ec	Corrected docs for XML::Parser.	2014-04-26 12:57:35 +02:00
Yorick Peterse	f53fe4ed7c	Reset the lexer when resetting the parser. Also removed the unused @lines instance variable.	2014-04-25 00:15:24 +02:00
Yorick Peterse	83ff0e6656	Various small parser cleanups.	2014-04-25 00:07:53 +02:00
Yorick Peterse	ecf6851711	Revert "Move linking of child nodes to a dedicated mixin." This doesn't actually make things any easier. It also introduces a weirdly named mixin. This reverts commit `0968465f0c`.	2014-04-24 21:16:31 +02:00
Yorick Peterse	0968465f0c	Move linking of child nodes to a dedicated mixin.	2014-04-24 09:43:50 +02:00
Yorick Peterse	08d412da7e	First shot at removing the AST layer. The AST layer is being removed because it doesn't really serve a useful purpose. In particular when creating a streaming parser the AST nodes would only introduce extra overhead. As a result of this the parser will instead emit a DOM tree directly instead of first emitting an AST.	2014-04-21 23:05:39 +02:00
Yorick Peterse	9ee9ec14cb	Lexer: only pop elements when needed.	2014-04-19 01:10:32 +02:00
Yorick Peterse	54e6650338	Don't use define_method in the lexer. Profiling showed that calls to methods defined using `define_method` are really, really slow. Before this commit the lexer would process 3000-4000 lines per second. With this commit that has been increased to around 10 000 lines per second. Thanks to @headius for mentioning the (potential) overhead of define_method.	2014-04-17 19:08:26 +02:00
Yorick Peterse	d9fa4b7c45	Lex input as a sequence of bytes. Instead of lexing the input as a raw String or as a set of codepoints it's treated as a sequence of bytes. This removes the need of String#[] (replaced by String#byteslice) which in turn reduces the amount of memory needed and speeds up the lexing time. Thanks to @headius and @apeiros for suggesting this and rubber ducking along!	2014-04-17 17:45:05 +02:00
Yorick Peterse	70516b7447	Yield tokens in the lexer and parser. After some digging I found out that Racc has a method called `yyparse`. Using this method (and a custom callback method) you can `yield` tokens as a form of input. This makes it a lot easier to feed tokens as a stream from the lexer. Sadly the current performance of the lexer is still total garbage. Most of the memory usage also comes from using String#unpack, especially on large XML inputs (e.g. 100 MB of XML). It looks like the resulting memory usage is about 10x the input size. One option might be some kind of wrapper around String. This wrapper would have a sliding window of, say, 1024 bytes. When you create it the first 1024 bytes of the input would be unpacked. When seeking through the input this window would move forward. In theory this means that you'd only end up with having only 1024 Fixnum instances around at any given time instead of "a very big number". I have to test how efficient this is in practise.	2014-04-17 00:39:41 +02:00
Yorick Peterse	25edd2de00	Use a Set for storing void element names.	2014-04-10 12:28:47 +02:00
Yorick Peterse	b96f7c4852	Lex attributes with namespaces. These are lexed as just the name instead of two separate tokens.	2014-04-10 11:01:49 +02:00
Yorick Peterse	c974b96b88	Truncate lines in parser errors. The offending lines of code displayed in the error message are truncated to 80 characters. This should make reading the error messages less of a pain when dealing with very long lines of HTML/XML.	2014-04-10 10:08:51 +02:00
Yorick Peterse	8237d5791d	Stream tokens when lexing. Instead of returning the tokens as a whole they are now streamed using XML::Lexer#advance. This method returns the next token upon every call. It uses a small buffer in case a particular block of text results in multiple tokens.	2014-04-09 22:08:13 +02:00
Yorick Peterse	e9bb97d261	First steps towards making the lexer stream tokens	2014-04-09 19:32:06 +02:00
Yorick Peterse	cb74c7edf9	Specs for XML parser errors.	2014-04-07 21:31:36 +02:00
Yorick Peterse	54ef125637	Basic docs for everything under Oga::XML.	2014-04-04 17:48:36 +02:00
Yorick Peterse	13a9228563	Properly indent doctype/XML decl inspect values.	2014-04-04 11:13:39 +02:00
Yorick Peterse	37a12722cb	Rough setup for a custom #inspect format. This format is a lot more readable than the default Ruby #inspect format (mostly due to not including previous/next/parent nodes).	2014-04-04 00:41:29 +02:00
Yorick Peterse	a2c525dd7c	Insert newlines after XML dec/doctypes.	2014-04-03 23:04:21 +02:00
Yorick Peterse	230fafa2d3	Document should not inherit from Node. A document is not an XML node on itself. If logic has to be shared between the Document and the Node class I'll resort to using mixins for this.	2014-04-03 22:45:40 +02:00
Yorick Peterse	c077988dd6	Tree building of doctypes.	2014-04-03 22:44:00 +02:00
Yorick Peterse	81b1155af3	Lex/parse doctype names separately.	2014-04-03 21:59:57 +02:00
Yorick Peterse	8185656c1e	Fixed typ.	2014-04-03 21:41:31 +02:00
Yorick Peterse	30c01a5aee	Tests for XML::TreeBuilder#handler_missing.	2014-04-03 09:43:30 +02:00
Yorick Peterse	bdb76cefc5	Dedicated handling of XML declaration nodes.	2014-04-02 22:30:45 +02:00
Yorick Peterse	d6c0a1f3f3	Lex/parser XML declaration attributes.	2014-04-02 22:01:17 +02:00
Yorick Peterse	f99c13b516	Tests + docs for the TreeBuilder class.	2014-03-28 17:11:54 +01:00
Yorick Peterse	6d866523b8	Renamed XML::Builder to XML::TreeBuilder.	2014-03-28 16:37:37 +01:00
Yorick Peterse	e141c084f9	Dedicated DOM builder class for CDATA tags.	2014-03-28 09:27:53 +01:00
Yorick Peterse	2b250bbf42	Rough DOM building setup.	2014-03-28 08:59:48 +01:00
Yorick Peterse	6ae52c1b12	Initial rough sketches for the DOM API.	2014-03-26 18:12:00 +01:00
Yorick Peterse	4a48647d1e	Removed generated lexer/parser. I am a dumbass.	2014-03-25 21:47:40 +01:00
Yorick Peterse	fb626278a8	Re-wrapped comments in the XML lexer.	2014-03-25 10:12:39 +01:00
Yorick Peterse	8ebd72158c	Renamed XML::Lexer#t to #emit().	2014-03-25 09:42:52 +01:00
Yorick Peterse	79818eb349	Added a convenience class for parsing HTML. This removes the need for users having to set the `:html` option themselves.	2014-03-25 09:40:24 +01:00
Yorick Peterse	eae13d21ed	Namespaced the lexer/parser under Oga::XML. With the upcoming XPath and CSS selector lexers/parsers it will be confusing to keep these in the root namespace.	2014-03-25 09:34:38 +01:00
Yorick Peterse	2259061c89	Don't require the 2nd Lexer#add_token argument.	2014-03-24 21:35:47 +01:00
Yorick Peterse	641c54261e	Simplified lexer output for comments.	2014-03-24 21:34:30 +01:00
Yorick Peterse	eaf1669b07	Simplified lexer output for CDATA tags.	2014-03-24 21:33:05 +01:00
Yorick Peterse	470be5a839	Simplified the lexer output for doctypes.	2014-03-24 21:32:16 +01:00
Yorick Peterse	ac775918ee	Lexing/parsing of XML declaration tags. This closes #12.	2014-03-24 21:30:19 +01:00
Yorick Peterse	b695ecf0df	Renamed element lexer tags. T_ELEM_OPEN has been renamed to T_ELEM_START, T_ELEM_CLOSE has been renamed to T_ELEM_END. This keeps the token names consistent with the other ones (e.g. T_COMMENT_START).	2014-03-24 20:32:43 +01:00
Yorick Peterse	52abc9d29e	Basic documentation for Oga::Parser.	2014-03-23 21:29:57 +01:00
Yorick Peterse	19c1d66287	Use String#unpack instead of String#codepoints. The latter returns an Enumerable which on Ruby 1.9.3 doesn't have #length available. Besides this it's better to just return an Array since we'll iterate over every character anyway.	2014-03-23 21:21:27 +01:00
Yorick Peterse	a2452b6371	Use codepoints instead of chars in the lexer. Grand wizard overlord @whitequark recommended this as it will bypass the need for creating individual String instance for every character (at least not until needed). This becomes noticable on large inputs (e.g. 100 MB of XML). Previously these would result in the kernel OOM killing the process. Using codepoints memory increase by a "mere" 1-1,5 GB.	2014-03-23 20:20:07 +01:00
Yorick Peterse	cdf5f1d541	Improve lexer performance by 20x or so. This was a rather interesting turn of events. As it turned out the Ragel generated lexer was extremely slow on large inputs. For example, lexing benchmark/fixtures/hrs.html took around 10 seconds according to the benchmark benchmark/lexer/bench_html_time.rb: Rehearsal -------------------------------------------------------- lex HTML 10.870000 0.000000 10.870000 ( 10.877920) ---------------------------------------------- total: 10.870000sec user system total real lex HTML 10.440000 0.010000 10.450000 ( 10.449500) The corresponding benchmark-ips benchmark (bench_html.rb) presented the following results: Calculating ------------------------------------- lex HTML 1 i/100ms ------------------------------------------------- lex HTML 0.1 (±0.0%) i/s - 1 in 10.472534s 10 seconds for around 165 KB of HTML was not acceptable. I spent a good time profiling things, even submitting a patch to Ragel (https://github.com/athurston/ragel/pull/1). At some point I decided to give a pure C lexer + FFI bindings a try (so it would also work on JRuby). Trying to write C reminded me why I didn't want to do it in C in the first place. Around 2AM I gave up and went to brush my teeth and head to bed. Then, a miracle happened. More precisely, I actually gave my brain some time to think away from the computer. I said to myself: What if I feed Ragel an Array of characters instead of an entire String? That way I bypass String#[] being expensive without having to change all of Ragel or use a different language. The results of this change are rather interesting. With these changes the benchmark bench_html_time.rb now gives back the following: Rehearsal -------------------------------------------------------- lex HTML 0.550000 0.000000 0.550000 ( 0.550649) ----------------------------------------------- total: 0.550000sec user system total real lex HTML 0.520000 0.000000 0.520000 ( 0.520713) The benchmark bench_html.rb in turn gives back this: Calculating ------------------------------------- lex HTML 1 i/100ms ------------------------------------------------- lex HTML 2.0 (±0.0%) i/s - 10 in 5.120905s According to both benchmarks we now have a speedup of about 20 times without having to make any further changes to Ragel or the lexer itself. I love it when a plan comes together.	2014-03-23 12:46:22 +01:00
Yorick Peterse	0e9d9b844c	Removed duplicate start_element rule.	2014-03-21 18:54:47 +01:00
Yorick Peterse	56ed9e949c	Use index based buffering for strings. This uses the same system as for T_TEXT nodes.	2014-03-21 17:45:40 +01:00
Yorick Peterse	9fa694ad4f	Use index based buffers for text nodes. Instead of appending single characters to a String buffer the lexer now uses a start and end position to figure out what the buffer is. This is a lot faster than constantly appending to a String.	2014-03-21 17:32:07 +01:00
Yorick Peterse	55f116124c	Fix for showing lines in parser errors.	2014-03-21 00:16:20 +01:00
Yorick Peterse	7749f4abce	Corrected a comment in the parser.	2014-03-21 00:10:20 +01:00
Yorick Peterse	a20ec0000a	Show up to 5 surrounding lines in parser errors.	2014-03-20 23:40:25 +01:00
Yorick Peterse	91fb7523fd	Lex open tags with newlines in them.	2014-03-20 23:39:29 +01:00
Yorick Peterse	ba17996bfc	Fancier error messages for the parser. The error messages of the parser now contain surrounding lines of code instead of only the offending line of code. This should make debugging a bit easier. Line numbers are also shown for each line.	2014-03-20 23:30:24 +01:00
Yorick Peterse	74bc11a239	Rip out column counting. This makes both the lexer and parser quite a bit easier to use. Counting column numbers isn't also really needed when parsing XML/HTML.	2014-03-20 19:44:28 +01:00
Yorick Peterse	70a39042e7	Removed useless rules from the parser.	2014-03-20 18:58:32 +01:00
Yorick Peterse	03774f2788	Documented the lexer.	2014-03-19 22:05:57 +01:00
Yorick Peterse	f1fcdfbacb	Cleaned up the Ragel bits of the lexer. This removes some of the complexity that existed before (e.g. too many state machines) and fixes a bunch of problems with nested data.	2014-03-19 21:44:10 +01:00
Yorick Peterse	7271e74396	Revert "Compacter parser AST." Although this AST is compacter it will result in conflicts between (text), (attributes) and (attribute) nodes in regular XML documents. This is due to XML allowing elements with these names (unlike in HTML). This reverts commit `8898d08831`.	2014-03-18 18:55:16 +01:00
Yorick Peterse	9975c9c430	Removed the emit_text_buffer Ragel action.	2014-03-17 21:49:49 +01:00
Yorick Peterse	274ab359ba	Don't use separate tokens/nodes for newlines. Newlines are now lexed together with regular text. The line numbers are advanced based on the amount of "\n" sequences in a text buffer.	2014-03-17 21:26:21 +01:00
Yorick Peterse	8898d08831	Compacter parser AST. The AST no longer uses the generic `element` type for element nodes but instead changes the type based on the element type. That is, a <p> element now results in an (p) node, <link> in (link), etc.	2014-03-17 21:03:54 +01:00
Yorick Peterse	cb75edc30d	Basic support for lexing/parsing HTML5. This will need a bunch of extra tests before I'll consider closing #7.	2014-03-16 23:42:24 +01:00
Yorick Peterse	ce8bbdb64a	Parsing support for multiple nested nodes.	2014-03-15 20:19:54 +01:00
Yorick Peterse	05ee3c13c9	Parsing support for nested element/text nodes.	2014-03-14 00:44:11 +01:00
Yorick Peterse	6b2f682c5c	Tests for lexing a basic HTML document. This also comes with some changes to the lexer so that it advances column/line numbers correctly.	2014-03-13 23:55:18 +01:00
Yorick Peterse	34f8779c94	Lexing of bare regular text. This is currently a bit of a hack but at least we're slowly getting there.	2014-03-13 00:42:12 +01:00
Yorick Peterse	2fbca93ae8	Supported for parsing nested elements.	2014-03-12 23:13:28 +01:00
Yorick Peterse	8cfa81aed9	Basic support for parsing elements. This includes support for elements with namespaces and attributes. Nested elements are not yet supported.	2014-03-12 23:02:54 +01:00
Yorick Peterse	5ce515d224	Small line wrapping change in the lexer.	2014-03-12 22:42:13 +01:00
Yorick Peterse	98b3443e7f	Lexing of element attributes without values.	2014-03-12 22:41:17 +01:00
Yorick Peterse	ed9d8c05a2	Added support for parsing comments.	2014-03-12 22:20:12 +01:00
Yorick Peterse	0a396043f8	Support for parsing CDATA tags.	2014-03-11 22:22:02 +01:00
Yorick Peterse	c9592856f0	Updated parsing of doctypes. The resulting nodes now separate the type, public and system IDs in to separate string values.	2014-03-11 22:08:21 +01:00
Yorick Peterse	c07edc767b	Updated the gitignore entry for the parser.	2014-03-11 22:03:02 +01:00
Yorick Peterse	8ce76be050	Moved the parser class to Oga::Parser. Oga will use the same parser for XML and HTML so it doesn't make sense to separate the two into different namespaces (at least for now).	2014-03-11 22:01:50 +01:00
Yorick Peterse	77b40d2e81	Use a separate machine for closing tags. This makes it easier to advance column numbers for whitespace as well as captuing and emitting tokens for the closing tag.	2014-03-11 21:55:36 +01:00
Yorick Peterse	eacd9b88cf	Reworked token generation for elements. This emits separate tokens for the start tag (T_ELEMENT_OPEN) and name (T_ELEMENT_NAME). This makes it easier to include the namespace of an element (T_ELEMENT_NS) in the output.	2014-03-10 23:50:39 +01:00
Yorick Peterse	cd53d5e426	Fixed advancing column numbers. In a bunch of cases the column number would not be increased correctly.	2014-03-07 23:54:56 +01:00
Yorick Peterse	a5a3b8db3f	Basic lexing of HTML tags. The current implementation is a bit messy. In particular the counting of column numbers is not entirely the way it should be. There are also some problems with nested tags/text that I still have to resolve.	2014-03-03 22:08:46 +01:00
Yorick Peterse	d9ef33e1f8	Lexing of comments. This fixes #4.	2014-02-28 23:27:23 +01:00
Yorick Peterse	92ae48f905	Use fcall + fret instead of fgoto. This removes the hardcoded return to the main machine.	2014-02-28 23:19:31 +01:00
Yorick Peterse	30d3e455d1	Use squote/dquote everywhere in the lexer.	2014-02-28 23:18:23 +01:00
Yorick Peterse	970ce27283	Cleanup of buffering text/strings. This removes the need to use \|\|= and such, which should speed things up a bit and keeps the code cleaner.	2014-02-28 23:16:01 +01:00
Yorick Peterse	ca6f422036	Lexing of doctypes. This comes with various structural changes to the lexer as I'm slowly starting to get the hang of Ragel. Ragel is a beast but damn it's an awesome piece of software. Note that the doctype public/system IDs are lexed as T_STRING. The parser will figure out whether a ID is a public or system ID based on the order. This fixes #1	2014-02-28 23:08:55 +01:00
Yorick Peterse	3c825afee0	Cleaned up lexer rules a bit. There's no benefit to adding variables for angle brackets and such, it's much easier to grok to just use them directly.	2014-02-28 20:09:13 +01:00
Yorick Peterse	2294bf19f4	Better lexing of CDATA tags. This means the lexer is now capable of lexing CDATA tags that contain text such as ]].	2014-02-28 20:05:12 +01:00
Yorick Peterse	6138945d53	Moved some of the CDATA docs around.	2014-02-28 00:04:44 +01:00
Yorick Peterse	4883ac7384	Lexing of CDATA tags.	2014-02-28 00:03:37 +01:00
Yorick Peterse	2c82f88f6c	Basic lexing + parsing of doctypes. We're doing these the lazy way. I can't be bothered writing patterns/rules for 4 different formats for something such as doctypes.	2014-02-27 01:27:51 +01:00
Yorick Peterse	91f416f035	Moved ending tags into their own racc rule.	2014-02-26 22:20:11 +01:00
Yorick Peterse	4f04fa0d30	Untrack Racc generated files. Yorick, you can stop being bad now.	2014-02-26 22:18:33 +01:00
Yorick Peterse	e764ba640a	Basic parser setup without tests. Who needs tests anyway!	2014-02-26 22:17:47 +01:00
Yorick Peterse	c4e0406ed9	Lexing of CDATA tags.	2014-02-26 22:01:07 +01:00
Yorick Peterse	0a336e76d3	Renamed T_EXCLAMATION to T_BANG. This is way easier to type.	2014-02-26 21:54:27 +01:00
Yorick Peterse	684eccd3e2	Lex dashes as T_DASH instead of T_TEXT.	2014-02-26 21:52:32 +01:00
Yorick Peterse	39bbe5afc4	Expanded lexer tag/attribute tests.	2014-02-26 21:48:46 +01:00
Yorick Peterse	d32888f803	Basic lexer setup/tests. Too lazy to do this the right way. ᕕ(ᐛ)ᕗ	2014-02-26 21:36:30 +01:00
Yorick Peterse	5755c325bd	Imported a half-assed lexer.	2014-02-26 19:54:11 +01:00
Yorick Peterse	702477ca28	Basic project layout.	2014-02-26 19:50:16 +01:00

... 4 5 6 7 8 ...

429 Commits